frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

https://prismml.com/news/bonsai-2-27b
48•JonSchneider•1h ago

Comments

kamranjon•52m ago
Love this for the folks with 16gb graphics cards - 3.8 27b has been incredible but not quite runnable on anything less than 32gb - will try loading this up on my 16gb intel b50 and see how it goes - not sure these quants can be accelerated by the XPU cores yet but maybe in time!
kadoban•25m ago
You can run the ~4 bit quant(s) on 24gb, if you're not _too_ picky on context size.

This will hopefully be better, though it'd be a _very_ surprising increase in performace at the size they say. Would love to see more about how it benchmarks.

spijdar•18m ago
I run Unsloth's UD-Q4_K_S on 20 GB of VRAM (RX 7900 XT) and I get ~90k tokens of context without quantizing KV cache. With 8-bit quantization, I get about a 134k token context window. That's with only one slot, but for me, it works pretty darn well, with 20-35 tok/s depending on how full that window is.
abraxas•51m ago
I'm not following the local mdoel scene too closely but this seems quite amazing. Is this able to be run on Apple silicon too?
kamranjon•50m ago
"Ternary Bonsai 2 27B reaches up to 143 tokens/second on NVIDIA GeForce RTX 5090 and 46.8 tokens/second on M5 Max. On an RTX 4090, Ternary Bonsai 2 27B consumes just 0.714 mWh/token, making it 40% more energy-efficient than an 8B model running in full-precision."
pizza234•25m ago
Their mention of the 5090 is bit odd, since on 32 GB GPUs, Q6 fits while having better quality. Very interesting model for 16 GB GPUs though!
Havoc•31m ago
Their first 27B bonsai was able to run on an iphone.
Aurornis•22m ago
These are small enough that you can run them entirely in the browser https://huggingface.co/spaces/webml-community/ternary-bonsai...

Remember to clear the downloaded weights afterward.

Like the last model, it's amazing they work as well as they do. Use it for any longer task and they fall apart spectacularly and in interesting ways.

outofpaper•15m ago
So you have some fun examples?
z2•18m ago
I'd love to see a Bonsai model start with a 100B+ parameter model and get that down to <30 GB. But maybe at that point we call it Topiary?
JonSchneider•8m ago
I'm hoping they release an 8B v2 based on the Qwen 3.8 series in the near future - that would give us a really powerful model that could be run directly on users phones.
simonw•6m ago
If you want to try out out the GGUFs from https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf#th... be aware that you need Prism's llama.cpp fork to get them to work, from https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-...

Astra for Law

https://openai.com/index/astra-for-law/
176•vertigoruntime•2h ago•164 comments

Bend – A language that blocks AI mistakes via proof, on CPU and GPU

https://bend-lang.com/
161•nicolas-siplis•1h ago•76 comments

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

https://prismml.com/news/bonsai-2-27b
58•JonSchneider•1h ago•13 comments

Hister: A private search engine for the pages you visit and the files you keep

https://github.com/asciimoo/hister
364•bookofjoe•5h ago•118 comments

I Hate You Microsoft

https://henriquenunez.eu/posts/you_did_it_again_ms/
91•henriquenunez•37m ago•30 comments

Wax motor

https://en.wikipedia.org/wiki/Wax_motor
149•mhb•1d ago•30 comments

Fujitsu launches made-in-Japan next-generation CPU FUJITSU-MONAKA

https://global.fujitsu/en-global/pr/news/2026/09/14-02
466•my123•2d ago•171 comments

Sex, AI, and the Apocalypse

https://www.iankduncan.com/personal/2026-09-16-sex-ai-and-the-apocalypse/
14•Anon84•1h ago•1 comments

Everybody's Lost Their Minds

https://www.netmeister.org/blog/everybodys-lost-their-minds.html
229•ibobev•2h ago•130 comments

CrowdSec Source Code Leak

https://www.crowdsec.net/blog/crowdsec-statement-source-code-exposure
113•eccgecko•6h ago•33 comments

Rate limits on GitLab.com are changing

https://about.gitlab.com/blog/rate-limit-change-2026/
136•darkwater•6h ago•102 comments

Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data

https://arxiv.org/abs/2609.18842
84•Betelbuddy•5h ago•25 comments

Flet 1.0 – Build cross-platform apps in Python

https://flet.dev/
10•absqueued•1h ago•3 comments

The American Religion of Self-Storage Facilities

https://www.newyorker.com/magazine/2026/09/21/the-american-religion-of-self-storage-facilities
164•pseudolus•9h ago•283 comments

How GLM built its own inference infrastructure

https://z.ai/blog/glm-built-its-inference-infrastructure
346•whiteros_e•13h ago•254 comments

Why I didn’t sign the Fields medallists’ letter

https://gowers.wordpress.com/2026/09/17/why-i-didnt-sign-the-fields-medallists-letter/
180•simianwords•13h ago•234 comments

Zettascale (YC S24) Is Hiring ASIC/FPGA Engineers to Build Chips for ASI

https://zscc.ai/careers?job_id=109821
1•el_al•5h ago

TSMC revealing details about next gen A14 node

https://iedm26.mapyourshow.com/8_0/sessions/session-details.cfm?scheduleid=331
68•osnium123•2d ago•25 comments

Towards Self-Driving Codebases

https://blog.detail.dev/posts/towards-self-driving-codebases/
86•wilhelmklopp•5h ago•71 comments

How do we prevent mathemathics from devolving into the Medieval Era of secrecy?

https://mathoverflow.net/questions/515260/how-do-we-prevent-mathematics-from-devolving-into-the-m...
49•jjgreen•2d ago•25 comments

Running Ubuntu on the Lenovo IdeaPad Duet

https://vhaudiquet.fr/blog/duet-ubuntu/
57•vhaudiquet•3d ago•17 comments

André Weil and the Hodge Conjecture

https://jiahao116.github.io/Articles/
23•caojiahao•5d ago•7 comments

How Uber Protects Against Retry Storms

https://www.uber.com/us/en/blog/protecting-against-retry-storms/
5•iscmt•1h ago•0 comments

One year of sponsored Servo development

https://servo.org/blog/2026/09/15/one-year-of-sponsorship/
335•AshleysBrain•14h ago•136 comments

GraphViz Pocket Reference – Make a Graph

https://graphs.grevian.org/graph/6322783643500544
6•vismit2000•2d ago•0 comments

Launch HN: Skillsync (YC W26) – AI chat sessions made portable across agents

36•cat-whisperer•5h ago•42 comments

CCC invites all model citizens to 40C3

https://events.ccc.de/en/2026/09/12/40c3-model-citizens/
309•antonly•14h ago•171 comments

Canto: A speech model built for the real world

https://wisprflow.ai/canto
21•sleepypandas•3h ago•9 comments

The Return of Sail Power: Cargo Ships Are Turning Back to the Wind

https://gcaptain.com/the-return-of-sail-power-cargo-ships-are-turning-back-to-the-wind/
170•gumby•21h ago•119 comments

Show HN: Share your AI Setup, Learn from others

https://mysetup.ai/
157•steveybrown•9h ago•81 comments