frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Running Kimi K3 on a M1 Max

https://github.com/gavamedia/deltafin
45•tito•59m ago

Comments

lostmsu•37m ago
under 0.02 tok/s
Azantys•32m ago
0.01 tk/s is unusable for anything, you would wait a whole day for just 1000 token of output, what is the point of projects like this?
tito•29m ago
I like seeing the latest and greatest model crammed into new systems to see how it fares. To deal with the speed, one person on reddit suggested using it in an email interface rather than a chat interface.
hugopuybareau•22m ago
Love the email idea
tito•17m ago
Having the right type of interface makes a huge difference.

It reminds me of when Willow Garage chose to name their bot the TurtleBot, because if they named it anything else, people would think it was fast and capable. But when they called it Turtle Bot, people just kind of liked it and were satisfied with what it did.

At the level of Kimi 3, I probably can code only about 1,000 good tokens per day, too. (thankfully coding isn't my job)

magicalhippo•16m ago
> one person on reddit suggested using it in an email interface rather than a chat interface

Kimi Pen Pal. Bring back lettets and postcards. Do OCR, and use one of those 3D printer-like pen plotters write the model output as a letter.

Challenge would be automating the opening and OCR preparation, and the folding and mailing of the return letter. But given it's done commercially it should be possible.

embedding-shape•14m ago
Email would indeed be fitting for K3 running on a M1 Mac, as it'd take days/weeks to receive a response, which matches with my real-world emailing experience pretty well.
SXX•22m ago
It is fun.

Also its answering the question of what gonna happen if you wake up tomorrow and datacenters are gone. Or internets are gone.

Some people on our globe live in countries with no internet whatsoever. Of course most of them dont have Macbook with 64GB RAM either, but it's much much easier to get than internet connection or rack of GB200.

SOTA LLMs are efficiently compression of all the knowkedge humanity has built. Having ability to run it at home to extract said knowledge is important no matter the speed.

winstonp•13m ago
16 tokens / s is not nothing.
ggm•11m ago
So you subscribe to the belief we won't in future find mentalism in other galaxies or solar systems which operate on mechanisms we don't understand and think v e r y s l o w w w w w w w l y ?

(note. I am not a believer in AGI)

"useful" is highly contextual. The clock of the long "now" is not useful in the sense you mean, to synchronise your wristwatch. I'm still glad it exists.

antirez•32m ago
SSD streaming on an M5 Max 128GB: https://x.com/antirez/status/2082136334160818528

Soon decent speed across two Mac Studios with 512GB of RAM.

tito•30m ago
wow cool! I like watching the new models come out and how they end up crammed in to run on local machines. I learned what mxfp4 is thanks to this latest Kimi release - although it sounds like it means that there's less room for compression in the model compared to others.
als0•30m ago
Says it requires a 2TB disk? Must it be internal NVMe?
tito•28m ago
The Github specifically mentions an option to stream it from an online host. It's extremely slow.
Fergusonb•15m ago
You can use an external drive if it's mounted as a writable volume. I would make sure it's fast, maybe thunderbolt 3/4/5 enclosure with a fast drive.
piterrro•24m ago
Will it fit on ESP32??
tito•23m ago
1 button Kimi morse code interface
embedding-shape•15m ago
Better questions, how many ESP32s would it take to reach 1 tok/s decoding speed with K3?
nlessard•19m ago
Anyone who knows the state of NVMe hardware more than me know if this would obliterate the lifespan of your drive? Seems like the biggest limitation to me (some people are probably fine with letting their Macs churn over the weekend).
addaon•17m ago
Reads are not generally life-limiting for flash. (Well, no more so than power-on time in general. You still have aging mechanisms like electromigration, but these are orders of magnitude slower than write-induced damage.)
nlessard•15m ago
Cool, thanks
tjwebbnorfolk•13m ago
> ~60–76 s/token

I don't know if I'd call this "running"

sermah•10m ago
Had the same feeling when I first saw min/km units in some (human) running context.

UPD: I know it's not the same at all, just the reversal of units that gets me

tito•10m ago
I commented similarly below, but as a terrible programmer, I probably perform about 1 minute per token too (at Kimi 3 level). It puts into context how I think about intelligence
brokencode•6m ago
Maybe in terms of code produces, but one token is only a fragment of a thought for an LLM.

It’d be like thinking as slowly as Ents talk to each other in Lord of the Rings.

mips_avatar•8m ago
Would be interesting to see how fast it would be on 4x mac studio 512gb machines.
acmnrs•5m ago
The title should probably be edited to specify "M1 Max" instead of "M1 Mac". You aren't running K3 on a base M1 anytime soon. Either way, still a very impressive project.
tito•4m ago
Done, Mac -> Max

OpenAI just open-sourced Codex Security

https://github.com/openai/codex-security
189•bakigul•1h ago•37 comments

Substack writers, you need a website

https://elizabethtai.com/2026/06/10/substack-writers-you-need-a-website/
348•speckx•5h ago•187 comments

Half-Life ported to Mac OS 9

https://mac-classic.com/news/half-life-ported-to-mac-os-9/
54•freediver•1h ago•20 comments

Steel Bank Common Lisp version 2.6.7

https://sbcl.org/all-news.html?2.6.7
169•tmtvl•5h ago•63 comments

Anthropic publishes a practical key-recovery attack on HAWK-256

https://github.com/anthropics/cryptography-research-demo
23•bakigul•1h ago•1 comments

Offer rates for tech jobs fell from 51% to 39% since 2015

https://www.interviewquery.com/p/tech-interview-offer-rates-lowest-12-years
7•racketracer•10m ago•0 comments

Kimi K3 Architecture Overview and Notes

https://sebastianraschka.com/blog/2026/kimi-k3-architecture-notes.html
236•ModelForge•6h ago•30 comments

The iPhone Upgrade Program is being replaced by Apple Upgrade

https://www.apple.com/shop/iphone/iphone-upgrade-program
105•lkurtz•4h ago•183 comments

Delayed Gratification – Proud to Be 'Last to Breaking News'

https://www.slow-journalism.com/
207•speerer•6h ago•113 comments

MCP 2026-07-28 Specification: transport going stateless

https://blog.modelcontextprotocol.io/posts/2026-07-28/
90•Eldodi•3h ago•28 comments

Running Kimi K3 on a M1 Max

https://github.com/gavamedia/deltafin
46•tito•59m ago•29 comments

The Fabled Flatbreads of Uzbekistan (2015)

https://www.aramcoworld.com/articles/2015/the-fabled-flatbreads-of-uzbekistan
55•jxub•4d ago•36 comments

Zig's Incremental Compilation Internals

https://mlugg.co.uk/posts/incremental-compilation-internals/
159•garyhtou•6h ago•121 comments

Show HN: How far do I have to go to run into 100k people?

https://imjasonh.github.io/playground/population-rays/
26•ImJasonH•5d ago•16 comments

Recursion is lying to you

https://blog.gaborkoos.com/posts/2026-05-09-Your-Recursion-Is-Lying-to-You/
25•theanonymousone•2h ago•31 comments

Show HN: I was tired of opening 2 tabs for every HN link, so I made a userscript

https://github.com/twalichiewicz/HNewhere
3•twalichiewicz•25m ago•0 comments

Interview with Boris Cherny [video]

https://www.youtube.com/watch?v=qyPCVqFUyDo
26•knighthacker•23h ago•14 comments

Discovering Cryptographic Weaknesses with Claude

https://www.anthropic.com/research/discovering-cryptographic-weaknesses
145•gslin•5h ago•76 comments

How Do I Profile eBPF Code?

https://naveensrinivasan.com/posts/2026-07-22-how-do-i-profile-ebpf-code/
102•snaveen•6h ago•6 comments

New HIV vaccine shows unprecedented success in preclinical study

https://www.lji.org/news-events/news/post/new-hiv-vaccine-shows-unprecedented-success-in-preclini...
510•codebyaditya•9h ago•227 comments

Show HN: XY – A Fast, composable, GPU-accelerated interactive plotting library

https://github.com/reflex-dev/xy
99•apetuskey•6h ago•38 comments

Toolcraft

https://toolcraft.sh
14•handfuloflight•1h ago•1 comments

Pacing the frontier

https://www.pacingthefrontier.com/
36•reducesuffering•2h ago•20 comments

Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

https://arxiv.org/abs/2510.26692
260•ronfriedhaber•11h ago•111 comments

Una GPS smart watch – Repairable, USB-C charging, developer-friendly

https://unawatch.com/
113•pimterry•7h ago•77 comments

Hulios: An eBPF-powered, transparent Tor gateway for Linux

https://github.com/ghaziwali/Hulios
22•ghaziwali•2h ago•2 comments

Harmony Explained: Progress Towards a Scientific Theory of Music (2012)

https://arxiv.org/abs/1202.4212
84•surprisetalk•7h ago•65 comments

Robotics development made dead simple (open source)

7•Ekami•2d ago•2 comments

Anthropeum – Where in the world, and when, does this human artifact belong?

https://anthropeum.com/
126•bookofjoe•7h ago•35 comments

WOFF 1.0: a milestone on W3C's journey of fonts on the web

https://www.w3.org/blog/2026/woff-1-0-a-milestone-on-w3cs-journey-of-fonts-on-the-web/
58•hn_acker•5h ago•4 comments