frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Run Kimi K3 using 29 GB of RAM at 0.50 tok/s

https://github.com/sqliteai/waste
35•marcobambini•4h ago

Comments

cjbprime•54m ago
Does it not use Metal, on macOS? Would it be faster if it did?
marcobambini•18m ago
We tried to use Metal, but for that specific project it was slower than just using NEON ARM optimizations. It is all documented in the docs.
pja•45m ago
That README hits all my “this is authored by an LLM” instincts. I presume the codebase is also written by an LLM?
gruez•39m ago
>Contributors

>...

>claude

You don't need to presume. If someone is so lazy that they tell claude to commit their code (ie. they're too lazy to run git commit themselves), the chances they reviewed the code is slim.

k8sToGo•32m ago
Why do you say lazy? maybe they are ok with people seeing it is claude?
danirod•30m ago
To be fair, I appreciate when they are so upfront about who wrote the code without requiring further heuristics, so I encourage this behavior.
bensyverson•11m ago
Yes, I do this all the time, and also check in the co-authored project plans which drove the commits. For a project that is transparently only possible due to agentic coding, I don't see any reason to conceal the methods.
simonw•28m ago
Honestly, Claude writes better commit messages than most people.

Personally I've mostly given in to letting it commit for me now, though I do occasionally take over and hand-write the messages if it's a particularly important concept and Claude's is too verbose.

Codex/GPT-x defaults to one-line commit messages, which are too short. Claude likes to write several paragraphs, which is usually too long.

If you tell it how to commit properly once per session it will stick with your standards for the rest of that session, and you can put that in AGENTS.md if you can be bothered to.

marcobambini•13m ago
I wrote tons of software, even a programming language by hand https://github.com/marcobambini/gravity.

I'm using my skills to orchestrate LLMs and agents, and I can write better code much faster. As developers, we can choose to adapt to new technologies or become extinct.

jpecar•19m ago
Where can this 1tb k3.waste be downloaded?
marcobambini•17m ago
It is not yet available, the only way is to download the official Kimi K3 model and then convert it:

# 1. preflight: reachable? how big? does it fit? tools/fetch_weights.sh --dest /Volumes/staging/k3 --dry-run

# 2. download — resumable, safe to kill, safe to re-run tools/fetch_weights.sh --dest /Volumes/staging/k3

# 3. convert into a container uv run --with torch --with safetensors python tools/convert.py \ --src /Volumes/staging/k3 \ --out ~/models/k3.waste --jobs 3

logicallee•9m ago
Interesting project. The headline number (29 GB of RAM) is for 4k context.

From what I've read elsewhere, Kimi K3 is quite verbose in its thinking. At the quoted rate, it would generate only a total of 1.8k tokens in 1 hour. Is that enough for it to get any thinking done and produce output on more complicated prompts?

Elevators

https://john.fun/elevators
411•Jrh0203•3h ago•138 comments

Big Food vs. the People

https://www.lighthousereports.com/investigation/big-food-vs-the-people/
97•jruohonen•2h ago•39 comments

Getting 25 Gbps Thunderbolt Ethernet on My Mac Studio

https://www.jeffgeerling.com/blog/2026/getting-25g-ethernet-mac-thunderbolt/
50•speckx•2h ago•34 comments

Algorithms on billion-scale graph using 10GB RAM: I love DataFusion

https://semyonsinchenko.github.io/ssinchenko/post/datafusion-graphs-cc-2/
42•speckx•2h ago•9 comments

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

https://artificialanalysis.ai/models/deepseek-v4-flash
442•theanonymousone•10h ago•232 comments

DeepSeek-V4-Flash Update

https://api-docs.deepseek.com/updates/
584•dnhkng•12h ago•282 comments

A GTK4 SSH-askpass in Zig

https://xn--gckvb8fzb.com/a-gtk4-ssh-askpass-in-zig/
37•surprisetalk•2h ago•9 comments

The most official water costs $120k a gallon

https://signoregalilei.com/2026/07/26/the-most-official-water-costs-120000-a-gallon/
51•surprisetalk•3h ago•21 comments

Miso (YC S16) is hiring for U.S. expansion

https://www.ycombinator.com/companies/miso/jobs/g2uAlMG-founding-business-lead-u-s-expansion
1•victorology•1h ago

Arch Linux disables AUR package adoption

https://lwn.net/Articles/1086489/
80•database64128•4h ago•53 comments

The great wealth transfer reality check

https://usa.visa.com/partner-with-us/visa-consulting-analytics/economic-insights/great-wealth-tra...
19•MarcoDewey•1h ago•5 comments

13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS

https://swe-rebench.com
23•ibragim_bad•2h ago•6 comments

The End of an Era

https://hughhowey.com/the-end-of-an-era/
323•harscoat•6h ago•363 comments

qm

https://github.com/yc-software/qm
3•tosh•19m ago•0 comments

How JPEG works: Interactively explore JPEG's lossy compression methods

https://cgjennings.ca/articles/jpeg-compression/
21•at1as•4d ago•1 comments

Run Kimi K3 using 29 GB of RAM at 0.50 tok/s

https://github.com/sqliteai/waste
35•marcobambini•4h ago•12 comments

The session you cannot take with you

https://earendil.com/posts/session-portability/
693•apitman•14h ago•199 comments

Show HN: BitBang – Reach machines behind NAT from a browser, no account

https://github.com/richlegrand/bitbang-cli
34•narragansett•3h ago•5 comments

The C ``Clockwise/Spiral Rule''

https://c-faq.com/decl/spiral.anderson.html
39•etrvic•4h ago•21 comments

The Art of Decision-Making (2019)

https://www.newyorker.com/magazine/2019/01/21/the-art-of-decision-making
19•EndXA•2h ago•2 comments

Anti-fraud tools can't keep pace with scammers exploiting cheap internet calling

https://broadbandbreakfast.com/how-to-fight-back-against-fraudulent-robocalls/
35•dredmorbius•4h ago•33 comments

Google fixed more Chrome bugs in June than over the past two years, thanks to AI

https://blog.google/security/chrome-stronger-with-every-update/
428•Garbage•10h ago•424 comments

Danube's record low levels force shutdown of Hungary's only nuclear plant

https://www.bbc.com/news/articles/cn0nqv05g0do
84•vrganj•10h ago•102 comments

IMAX vs. IMAX 70mm: The difference between these two cinema formats

https://www.engadget.com/2220571/differences-between-imax-70mm/
59•ksec•6d ago•82 comments

Show HN: What should the GUI for AI agents look like?

https://marbleos.com/demo
89•akbabu•13h ago•55 comments

Solving poker in custom WebGPU kernels

https://phulin.me/blog/poker/
53•patrickhulin•1d ago•14 comments

Investigating three real-world incidents in our cybersecurity evaluations

https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
207•surprisetalk•19h ago•164 comments

Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

https://www.bloomberg.com/news/articles/2026-07-31/moonshot-s-kimi-built-on-20-000-nvidia-chip-cl...
81•gk1•5h ago•53 comments

Show HN: Slope remade in HTML5 to load instantly on any browser, any device

https://hurtle.site/
6•novlrdotcom•1h ago•1 comments

Winding Down Artichoke Ruby

https://hyperbo.la/w/winding-down-artichoke-ruby/
47•ksec•6d ago•7 comments