frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: A new engine to run Kimi K3 on a laptop

3•marcobambini•59m ago
Kimi K3 has 2.78 trillion parameters and ships as 1.42 TB of weights. It clearly does not fit in the memory of a laptop. But K3 is a Mixture-of-Experts model. For each token, only a small fraction of its 896 experts per layer is activated. That changes the problem: the entire model does not need to be resident in RAM, as long as the weights required by each token can be reached quickly enough.

We built WASTE — the Weight-Aware Streaming Tensor Engine — to explore that idea.

WASTE keeps the dense, repeatedly used part of the model resident in memory, stores the routed experts in an NVMe-optimized container, and streams only the experts selected during inference. The remaining RAM is used as a bounded expert cache.

The current Kimi K3 container is 982 GiB. On a 64 GB MacBook Pro, WASTE runs the complete model at around 0.32–0.34 tokens per second, with a measured minimum memory requirement of approximately 29 GB at a 4K context.

That is obviously not interactive performance yet. But the result we found interesting is that it works at all: this is the full open-weights model, not a distillation, a pruned version, or a smaller model using the Kimi name.

The engine is written in C and has no BLAS, CUDA, ONNX, or Python dependency in the inference path. The same code can be used through the CLI, embedded as a library, or exposed through the included OpenAI-compatible server.

Correctness was the first constraint. Every layer was validated against a PyTorch reference, with final logits matching within 3.6e-06. The vision tower is supported as well and matches its reference within 2.3e-06.

The current bottleneck is understood: K3 needs roughly 17 GB of expert data per token, and more than half of the decode time is spent reading experts from disk. The engine is already operating close to the measured throughput limit of the laptop’s internal SSD. The next improvements therefore need to reduce the number of bytes read per token and increase useful expert reuse without pushing the operating system into paging.

K3 is deliberately the extreme case. The same engine runs Kimi-Linear 48B from a 19 GB container at 8.92 tokens per second with an 8 GB memory budget. The broader goal is to make models that are much larger than available RAM usable locally, without sending private data to an API and without requiring specialized accelerator hardware.

We have published the engine, container format, conversion tools, benchmarks, validation suite, and also the experiments that failed rather than quietly removing them.

Everything is fully open source. Feedback on the storage layout, quantization, caching strategy, direct I/O, portability, and potential optimizations would be very welcome. Contributions of any kind — code, benchmarks, testing on different hardware, documentation, bug reports, or new ideas — are more than appreciated.

Repo: https://github.com/sqliteai/waste

Comments

tito•52m ago
What processor in the MBP?

eBay reaches $56M settlement with e-com newsletter writers it terrorized in 2019

https://techcrunch.com/2026/07/28/ebay-reaches-56m-settlement-with-e-commerce-newsletter-writers-...
1•pseudolus•34s ago•0 comments

Google Gemini Distillation Service [archive link]

https://web.archive.org/web/20260728173925/https://docs.cloud.google.com/gemini-enterprise-agent-...
1•bluepeter•1m ago•0 comments

Wish Is My Command

https://masterbran.ch/posts/your-wish-is-my-command.html
1•grapemane•1m ago•0 comments

Ton 618

https://en.wikipedia.org/wiki/TON_618
1•ethanpil•2m ago•0 comments

Show HN: Kudory – Tailor your CV for each job application

https://kudory.com
1•acbeni•3m ago•0 comments

Keychron announces first open-source firmware for gaming mice

https://www.digitalfoundry.net/news/2026/07/keychron-announces-first-open-source-firmware-for-gam...
2•JLO64•4m ago•1 comments

An agent with a spreadsheet engine beats one without

https://grid.is/blog/an-agent-with-a-spreadsheet-engine-beats-one-without
1•mooreds•4m ago•0 comments

The last decade of front-end engineering

https://www.natemeyvis.com/the-last-decade-of-front-end-engineering/
1•Brajeshwar•6m ago•0 comments

Every Time I Hire a Linguist, Inference Costs Go Down

https://arxiv.org/abs/2607.25335
2•cwbuilds•6m ago•0 comments

AI companies are reportedly shredding books after using them to train AI models

https://www.tomshardware.com/tech-industry/artificial-intelligence/ai-companies-are-reportedly-sh...
1•Pandionic•7m ago•0 comments

Boss Profits from Your Excuses

https://www.youtube.com/watch?v=suVyMze_qO0
1•gillysuit•7m ago•1 comments

NSF partners with universities and industry on pilot 4-year PhD programs

https://www.nsf.gov/news/nsf-partners-universities-industry-pilot-initiative-four
2•yiyingzhang•7m ago•1 comments

Show HN: Open-sourced, non-AI newtab extension to read your X/Twitter bookmarks

https://chromewebstore.google.com/detail/twitter-x-bookmarks-on-ne/acpkgdfhoaalmnhjifhneghcgfnjkglo
1•iankit17•10m ago•0 comments

Existing industry processes are blueprints for AI workflows

https://abtdomain.com/blog/2026/07/from-fords-assembly-line-to-local-ai-pipelines/
1•ABTdomain•11m ago•0 comments

eBay and ex-executives to pay $55.7M to couple sent cockroaches

https://www.theguardian.com/technology/2026/jul/28/ebay-harassment-lawsuit-settlement
1•mitchbob•11m ago•0 comments

Show HN: Energy, carbon and water estimates for AI content, shown as ranges

https://aicontentfootprint.com/
1•daniloedu•11m ago•0 comments

The rise of fake online shopping platforms that let you pretend to buy things

https://www.fastcompany.com/91560432/dopamine-sites-fake-online-shopping-apps-let-you-pretend-to-...
1•austinallegro•11m ago•0 comments

Emacs Writing Machine

https://chainsawriot.com/postmannheim/2026/07/25/writeredeck.html
1•abnercoimbre•14m ago•0 comments

Newsmax, Meta Enter AI Content Partnership

https://news.bloomberglaw.com/artificial-intelligence/newsmax-meta-enter-ai-content-partnership
1•cdrnsf•16m ago•0 comments

CryptanalysisBench: Can LLMs Do Cryptanalysis?

https://arxiv.org/abs/2607.18538
1•zdw•17m ago•0 comments

Ultimate Packer for eXecutables

https://upx.github.io/
2•jcbhmr•17m ago•0 comments

Replit Design

https://replit.com/blog/introducing-replit-design
2•ai2027•18m ago•0 comments

DoorDash Air, Our In-House Drone Delivery Program

https://about.doordash.com/en-us/news/doordash-air
1•ChrisArchitect•18m ago•0 comments

Enterprises are building their AI sales tools in-house and buying the data layer

https://www.octavehq.com/post/enterprise-gtm-ai-2026
1•connor11528•18m ago•0 comments

Hacktoberfest 2025

https://hacktoberfest.com
1•kushagra1117•19m ago•0 comments

We Created a Free Password Manager for AI Agents

https://serendb.com/blog/three-reasons-we-created-seren-passwords-for-ai-agents-before-humans
1•lewistaariq•21m ago•0 comments

DoorDash unveils its own drone delivery service after FAA approval

https://www.engadget.com/2226069/doordash-drone-delivery-service-faa-approval/
3•bookofjoe•22m ago•0 comments

Show HN: Agentsnap – Snapshot testing for AI agents

https://github.com/iamfaham/AgentSnap
3•iamfaham•23m ago•0 comments

A requiem for Intel's Optane, which could have eased the RAM price crunch

https://www.theregister.com/storage/2026/07/29/a-requiem-for-optane-intels-kv-cache-killer-that-c...
6•kencausey•24m ago•1 comments

Show HN: Kedge – Full-stack cloud with forkable VM snapshots and global SQLite

https://kedge.dev/
9•wgjordan•25m ago•1 comments