frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Kimi-K3 Releases on HuggingFace 7/27

https://huggingface.co/moonshotai/Kimi-K3
183•nateb2022•2h ago•60 comments

PGSimCity - How PostgreSQL Works

https://nikolays.github.io/PGSimCity/
542•jonbaer•8h ago•56 comments

Show HN: Physically accurate black hole you can put in your room

https://blackhole.plav.in
287•aplavin•3d ago•85 comments

Scriptc by Vercel: TypeScript-to-Native compiler, no JavaScript engine in binary

https://github.com/vercel-labs/scriptc
147•maxloh•9h ago•78 comments

French firefighters face 'pyrocumulonimbus' for first time

https://www.france24.com/en/live-news/20260726-french-firefighters-face-pyrocumulonimbus-for-firs...
344•saaaaaam•14h ago•221 comments

Decker, a platform that builds on the legacy of Hypercard and classic macOS

https://beyondloom.com/decker/
302•tosh•14h ago•73 comments

US citizen charged after GrapheneOS phone wipes during airport search

https://www.techspot.com/news/113236-us-prosecutors-charge-atlanta-man-after-grapheneos-phone.html
688•eecc•10h ago•493 comments

We have proof automation now

https://www.imperialviolet.org/2026/07/26/zstd-lean.html
160•zdw•11h ago•47 comments

I wanted a clock that never needed setting. Things escalated

https://arstechnica.com/gadgets/2026/07/i-wanted-a-clock-that-never-needed-setting-things-escalated/
94•lee_ars•3d ago•93 comments

Htmx 4.0, the first JavaScript library to release exclusively on the Game Boy

https://swag.htmx.org/en-cad/products/htmx-4-the-game
435•rcy•20h ago•145 comments

Introduction to Data-Oriented Design [pdf]

https://www.gamedevs.org/uploads/introduction-to-data-oriented-design.pdf
165•tosh•14h ago•44 comments

Fonts In Use – Find out where a font is used

https://fontsinuse.com/
60•open_•8h ago•4 comments

Simulate cassette tape audio profiles using FFmpeg

https://github.com/AARomanov1985/Audio-Cassette-Simulation
124•xterminal•12h ago•50 comments

Design is compromise

https://stephango.com/design-is-compromise
245•ankitg12•16h ago•81 comments

Measuring developer productivity with the DX Core 4

https://getdx.com/research/measuring-developer-productivity-with-the-dx-core-4/
10•saikatsg•2d ago•6 comments

8086 Emulator Inside Scratch

https://turbowarp.org/1248315967?size=640x400
4•rickcarlino•4d ago•1 comments

Show HN: CheapSecurity – Lightweight, Self-Hosted CCTV for Linux SBCs

https://github.com/gmrandazzo/CheapSecurity
128•zeldone•16h ago•26 comments

The Usefulness of Useless Knowledge (1939) [pdf]

https://faculty.lsu.edu/kharms/files/flexner_1939.pdf
71•jxmorris12•3d ago•4 comments

Show HN: Reverse Minesweeper

https://sunflowersgame.com/
213•pompomsheep•19h ago•68 comments

How to write English prose (2023)

https://thelampmagazine.com/blog/how-to-write-english-prose
117•geneticdrifts•15h ago•56 comments

Go Analysis Framework: modular static analysis by go team

https://pkg.go.dev/golang.org/x/tools/go/analysis
204•AbuAssar•20h ago•64 comments

The Zen of Parallel Programming: The Posture of a Kernel

https://smolnero.com/posts/the-zen-of-parallel-programming-the-posture-of-a-kernel
14•edgar_ortega•4d ago•1 comments

Kill The Cookie Banner

https://killthecookiebanner.eu/
979•rapnie•20h ago•467 comments

The New AI Superpowers: Focus and Followthrough

https://www.rickmanelius.com/p/the-new-ai-superpowers-focus-and
197•mooreds•19h ago•63 comments

I learned PCB design, 3D printing and C just to listen to music

https://pentaton.app/blog/2026-07-12-introducing-pentaton-lp/
206•interfeco•3d ago•42 comments

The relay market powering token resellers and fraud

https://vectoral.com/blog/token-relay-market
189•mlenhard•17h ago•117 comments

Some more things about Django I've been enjoying

https://jvns.ca/blog/2026/07/21/more-nice-django-things/
158•surprisetalk•5d ago•96 comments

Teaching Kids Forth

https://gracefulliberty.com/articles/teaching-kids-forth/
85•rbanffy•11h ago•30 comments

How to Block Some of the Bots

https://nochan.net/b/Internet-Crap/20260606-How-To-Block-Some-Of-The-Bots/
100•Bender•14h ago•118 comments

History of John Backus's functional programming project (draft)

https://softwarepreservation.computerhistory.org/FP/
30•cwbuilds•2d ago•2 comments
Open in hackernews

Kimi-K3 Releases on HuggingFace 7/27

https://huggingface.co/moonshotai/Kimi-K3
174•nateb2022•2h ago

Comments

NitpickLawyer•2h ago
This will be interesting for a few reasons. First, depending on where the median pricing settles w/ 3rd party providers will tell us what it costs to serve a 3T model. Since it's going to be mxfp4 native, it'll take ~1.5TB of VRAM to host this, which is juuust at the limit of 8xb200s (but realistically you'll need 16x for context / throughput optimisation). Won't be cheap to host, but at least we should get some range of $/MTok for a 3T model. Then we'll be able to guesstimate if "labs are subsidising tokens on API pricing".

Also interesting to see what effort it will take to fine-tune this beast. The latest AISI benchmarks on cybersec place it above glm5.2, but still way way behind SotA closed models. Some fine-tuning might be needed here. Also, interesting to see if Cursor does another training round on it, to directly compare it w/ kimi2.6/2.7 fine-tunes (composer series) and grok4.5.

Also also, interesting to see if someone takes on distilling (proper distillation, w/ training the entire distribution) from this into smaller models. (dsv4-kimi should be really good, since dsv4 is very cheap to serve)

dist-epoch•1h ago
> if "labs are subsidising tokens on API pricing"

> SemiAnalysis estimates that Anthropic's current blended gross margin has risen to the mid-60% range, with the API business gross margin exceeding 80%

Of course, people will insist "they are lying", "why should we believe them, it's well known they subsidize API pricing", ...

https://newsletter.semianalysis.com/p/anthropic-3q26-profit-...

https://finance.biggo.com/news/02d45650-b569-4d12-b44d-8d6d8...

NitpickLawyer•58m ago
Agreed. My (somewhat educated) guess is that top labs have healthy margins on API pricing. But this release will add another 3rd party / clear of conflict datapoint in this estimation.
knollimar•35m ago
That numner blends in training or no?
woctordho•57m ago
Speaking of finetune, currently a common practice is LoRA over bnb 4-bit base model, but I think it's time to replace bnb with GGUF as the base model format. GGUF is actively supporting new model architectures and more aggressive quantizations.

I've made some proof of concept in https://github.com/woct0rdho/transformers5-qwen3.5-recipe . We can finetune Qwen3.5-35B-A3B in 16 GiB VRAM, and DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM, without CPU offload. This works well on unified memory machines like Strix Halo.

Even so, larger models like Kimi-K3 still require multiple GPUs and nodes, and there are a lot more to do compare to single-GPU training.

vb-8448•26m ago
> Then we'll be able to guesstimate if "labs are subsidising tokens on API pricing".

No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models.

Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?)

amelius•21m ago
> Without training cost you can infer only the marginal cost of serving this kind of models.

Which is by far the most interesting number of the two.

> Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?)

If you get close in output quality, then does that matter?

walrus01•24m ago
It will be very interesting to see what kind of 'slow' performance people get from running it on a no GPU, but tons of RAM server (like a dual or quad socket xeon with 1.5 to 3TB of RAM). For the purpose of giving it longer duration tasks to generate a piece of something and come back and check on what it has done in 4 or 6 hours. Even if the output is like 5-6 tok/s, that might be usable for some purposes.

Huge price difference in what you can do with buying a used 4U rackmount server and putting 3TB of RAM in it (64GB DIMMs x quantity 32 in a quad socket xeon, you can see some benchmark prices on eBay for sets of 16 or 32 matched 64GB ECC DIMMs) for <$30,000, vs the cost of trying to run it on real GPU hardware.

Now obviously, as of the time I write this, the full precision hasn't been released nor has anyone like unsloth run it through quantization yet to produce a "Q8" or "Q8-XL" variant of it. But I think it's going to need more than 1536GB of RAM, with a usable and large amount of context, more like 2TB and preferably 2.5 to 3TB.

I also predict that people who try to run it in Q4 and Q6 will get the worst of both worlds, less precision/lost knowledge but also not reliable output that comes out too slow. In my personal opinion if I'm going to deal with something that is smart but slow and running on limited budget hardware, I need it to be Q8.

sandworm101•21m ago
Those old LTT videos of high core-count threadrippers running GPU benchmarks become more relevant each day.
walrus01•17m ago
The performance bottleneck is not really so much the number of cores or processing power in each core, but the memory bus bandwidth to/from the CPU. I have an older dual socket xeon server here which is a CPU-only LLM test machine with 256GB of RAM and the actual CPU stress is not much, I can even quantify this by how little it spins up the CPU fans to meet thermal load (the CPUs are operating at nowhere near their 180W per socket max capacity, compared to like, crunching prime numbers or running cpuburn).

But the memory bus speed is fully committed when generating tokens or thinking.

woadwarrior01•9m ago
Since the model is natively MXFP4, I think it'll be even more interesting on the hardware front. It'll comfortably fit on a 8x AMD MI355X node. I suspect that'll drive token prices down, further.
Der_Einzige•7m ago
Anyone who thinks that the labs are not profitable on per token API pricing is delusional and hilariously wrong.
m00dy•1h ago
There’s going to be a lot of competition around this model. Let’s see how low AI providers are willing to push prices.
Iolaum•1h ago
As long as they are transparent about what quant they serve the model and any other optimization they do that also affects performance of inferred tokens.
torginus•59m ago
I think the results might be underwhelming - AI providers need to turn a profit and can't subsidize, and they're working off of the commodity hardware everyone does.

I wouldn't be surprised if they started offering potentiall bad quantizations with much reduced capability at lower prices (without telling the users, of course)

wwwhizz•1h ago
That would be 7/27.
ra•1h ago
or 27/7 for the rest of the world
LeonM•52m ago
No, 27-7 for the rest of the world.

The separator is often the only way to distinguish American notation from ISO, so please use a dash for dd-mm-yy and a forward slash for mm/dd/yy

letier•49m ago
Or 27.7. in some other places.
casparvitch•48m ago
Have never seen 27-7, as someone in rest-of-world
kuboble•43m ago
Nope,

There are quite some countries around the world using d/m/y

https://en.wikipedia.org/wiki/List_of_date_formats_by_countr...

Algeria, Belgium, Brazil, Chile...

distances•38m ago
CodeCompost•1h ago
Why is there a countdown?
broodbucket•1h ago
You're not having a party?
InsideOutSanta•39m ago
I think it's shameful that Moonshot isn't providing us with party kits like Microsoft did with the Windows 7 Launch Party kit. How am I supposed to properly celebrate this without fun Kimi-themed quizzes for my guests?
mythz•1h ago
It's a release party
davidkunz•1h ago
This is historic. For the first time, an open-weights LLM is right at the top.

We won't be able to run this ourselves, but many providers can.

embedding-shape•32m ago
> For the first time, an open-weights LLM is right at the top.

Hmm, not quite true, I think that honor, for better or worse, goes to OpenAI. When they released GPT2 (or GPT1 for that matter) is was quite literally the SOTA in the ecosystem when it was released.

maelito•1h ago
Did someone run censorship and political bias tests on this ? Must be interesting.
walrus01•19m ago
As a completely one person, single sample anecdote, the 'heretic' uncensored Q8 GGUF variants several people have published of Qwen 3.5-122, 3.6-27B and 3.6-35B-A3B will very happily discuss just about any controversial topic that the CCP hates. Including lots of things that would get you thrown into prison if you published them in Mandarin on the domestic Chinese internet.

https://github.com/p-e-w/heretic

As a side note on this, if you see the reference in the screenshot in the link above to the harmful behaviors prompt set, these are all in English:

https://huggingface.co/datasets/mlabonne/harmful_behaviors

You could likely further de-censor a model by having a set of 'test' prompts in native Mandarin, Cantonese or really just about any other language. I don't speak any Chinese languages so I don't know if the published 'heretic' GGUF files some people have been throwing around will cooperate, or refuse, if you ask it in Mandarin for how to build a meth lab or precursors for semtex.

marvinLuck•1h ago
The weightings should be released on July 27.
sreekanth850•53m ago
how feasible its will be to run on modal or deepinfra? anyone here tried and tested such large models running?
throwaw12•52m ago
Strange communists, giving away such an expensive model to the public.

On the other note, can't wait to see 1bit quantisation soon and how it performs in benchmarks, if it performs really well in benchmarks, would be very good news for GPU hosting providers, to offer "Opus 4.5 level model at the cost of Haiku 4.5"

pmg1991•47m ago
Hoping no issues on Huggingface due to download rush.
someguyornotidk•20m ago
For huge models like these, the only reasonable way to host them is via torrents. I don't understand why hf doesn't offer this as an option.

Linux distributions got this right: Offer both HTTP and Torrents. Let the user decide.

gorgmah•47m ago
We already know that competition brought GLM 5.2 prices down roughly 45% since its release on June 16th (1.5 months ago), and the price downward slope is probably still going (I've been checking regularly and new providers keep fighting on price, I don't think prices have settled yet). For reference : https://openrouter.ai/z-ai/glm-5.2#providers

I saw arguments like "Providers cannot price less than their costs" in other comments. In economics, it's generally admitted that they shouldn't price less than their marginal costs, i.e. in their case roughly the cost of electricity, since a lot of these datacenters are not at capacity in terms of graphics cards usage (speculation since it's very easy to rent a GC for a couple hours on some providers). My guess is that someone will be selling tokens at less than electricity + depreciation of GCs soon, since there's a lot of competition and "smaller" data centers have overcapacity? This is speculation, correct me if I'm wrong

markasoftware•35m ago
Press x to doubt on the 45% number. The cheaper providers on open router are fp4 vs fp8 for official zai. There are some cheap fp8 ones (like novita) but the ui makes it seem like it's a temporary promotion, with their normal prices being almost equal to official zai (idk much about open router so not really sure what's going on with these discounts)
rs38•36m ago
is there a realistic way to distill 2 consumer hardware friendly models with max ~200B and ~20B? Qwen did it, but would it be possible for 3rd parties (unsloth etc)?
embedding-shape•34m ago
Yeah, why not. Toughest part is running the hardware so you can create the traces for downstream training, but once over that hump, nothing would stop you from doing that no.
colortiles•36m ago
This looks really promising. Excited to see where this goes. Looking forward to trying it out!
minimaxir•34m ago
...does Hugging Face have enough bandwidth to let people download en masse however much file size a 2 trillion parameters model is?
KronisLV•29m ago
I feel like most hardware to run LLMs on is shaped wrong for individuals.

It's either having a model struggling along with like 5-10 tokens per second on unified memory, or data center cards with hundreds of GB of VRAM consuming more than a kW of power. It doesn't seem like there's prosumer GPUs with like 180W-250W TDP and 128 GB or 256 GB of VRAM (one can dream). Then bifurcation and even just two of those cards would be kinda useful (albeit NVLink or equivalent would need to be commonplace).

Obviously nobody is running Kimi K3 locally without an insanely beefy homelab and lots of money to burn, but running GLM 5.2 would be cool at like ~100 tokens per second for a single session and maybe ~60 tokens per second with N subagents.

How unfortunate.

walrus01•10m ago
I have found that the "mostly didn't lose anything" Q8 large models that I want to run are all too large to run on the "only $3995!" 128GB max RAM systems that some people are buying, and definitely won't fit with any usable amount of context. Things like Qwen 3.5 122B Q8 or deepseek v4 flash Q8, or Laguna S 2.1 Q8 need 170-190GB of RAM including full context, which fits on a 256GB RAM dual socket workstation or rackmount server (sans GPU).

Copy and paste below from my notes and reported memory consumption with latest llama-server, assuming use of "--no-mmap" to load the entire thing into RAM at the time that llama-server launches.

DeepSeek-V4-Flash-UD-Q4_K_XL via unsloth 145GB on disk GGUF 0.03.323.204 I common_params_fit_impl: projected to use 178175 MiB of host memory

DeepSeek-V4-Flash-UD-Q8_K_XL via unsloth 151GB on disk GGUF 0.02.215.885 I common_params_fit_impl: projected to use 184636 MiB of host memory

Laguna-S-2.1-UD-Q8_K_X via unsloth 120GB on disk 0.01.616.119 I common_params_fit_impl: projected to use 172860 MiB of host memory

Qwen3.5-122B-A10B-UD-Q8_K_XL via unsloth 160GB on disk GGUF 165GB RAM use on launch, fresh context 0.04.976.905 I common_params_fit_impl: projected to use 170038 MiB of host memory

hellajack3d•11m ago
So... Now we give huggingface the hug of death - right? ;)
I've never seen dd-mm-yy. It's usually dd.mm.yy, dd.mm.yyyy, or yyyy-mm-dd, with some dd/mm/yy sprinkled in for general confusion.
ashwoods•24m ago
Not really. Spain's traditional format is dd/mm/yyyy with slashes. This applies for a good chunk of Europe. Germany/Austria uses dot, I think nordic countries embraced the dash. While you might see more adoption in offical/digital contexts, I just double checked a few popular spanish websites, all slashes.
pavo-etc•17m ago
This is so confidently wrong it's funny. In Australia dd/mm/yy is the default.
psychoslave•10m ago
Most life forms don’t use any calendar actually.
broodbucket•1h ago
20270727 if we're improving dates :)
m00dy•59m ago
Looks like another LeetCode problem about checking for anagrams.
c7b•58m ago
270727 if we want to save tokens :)
darkwater•47m ago
Y2K100 bug for you, then ;)
fragmede•55m ago
ISO 8601 ftw
linzhangrun•50m ago
China uses YYYY/MM/DD, which is logical.
_zoltan_•46m ago
the only logical format.

signed: a hungarian :)

JSR_FDED•27m ago
lpszReleaseDate ;-)
egeozcan•22m ago
For me, the only format that doesn't make sense is the MM/DD/YYYY, together with its rarely seen worse sibling, MM/DD/YY (07/27/26).