Ternary Bonsai: Top Intelligence at 1.58 Bits

41•nnx•2d ago

Comments

wmf•1h ago

Yet again they're comparing against unquantized versions of other models. They would probably still win but by a much smaller size margin.

Dumbledumb•54m ago

Wouldnt the margin be higher? All other models being moved from unquantized to quantized would lower their performance, while bonsai stays. I get what you see if it was in regards to score/modelsize, but not for absolute performance

mchusma•1h ago

Ever since I saw the first one of these one-bit models made by Microsoft, I thought this was a fascinating route. I assume that in practice, this is less helpful than it seems, just because there's every economic incentive in the world for the big AI labs to produce small, powerful, fast models. None of them seem to be using this technique, so it's interesting, but I suspect it's not quite working.

I also have yet to see any of these at a larger scale. For example, can you try one of these at 100 billion parameters?

yodon•1h ago

So excited to see this - the big advantage of 1.58 bits is there are no multiplications at inference time, so you can run them on radically simpler and cheaper hardware.

Animats•44m ago

At 4 bits, you could just have a hard-wired table lookup. Two 4 bit values in, 256 entry table. You can have saturating arithmetic and a post-processing function for free. Somebody must be building hardware like that.

Animats•46m ago

This makes sense. The 1-bit model implies needing 2x as many neurons, because you need an extra level to invert. But the ternary model still has a sign, just really low resolution.

(I've been reading the MMLU-Redux questions for electrical engineering. They're very funny. Fifty years ago they might have been relevant. The references to the Intel 8085 date this to the mid-1970s. Moving coil meters were still a big thing back then. Ward-Leonard drives still drove some elevators and naval guns. This is supposed to be the hand-curated version of the questions. Where do they get this stuff? Old exams?)

[1] https://github.com/aryopg/mmlu-redux/blob/main/outputs/multi...

ericb•40m ago

This is pretty cool! I would love to see an even larger models shrunk down.

If you got that into a couple gigs--what could you stuff into 20 gigs?

armanj•13m ago

I did a quick benchmark & compared it with Qwen3.5: https://github.com/ArmanJR/PrismML-Bonsai-vs-Qwen3.5-Benchma...

in my results, accuracy-wise Ternary-Bonsai-8B is on par with Qwen3.5-4B. But in accuracy-per-byte, bonsai is the clear winner:

=> Ternary-Bonsai-1.7B achieved 65.1% from 462 MiB, beating Qwen3.5-0.8B by 12 points while being ~5% smaller on disk. => Ternary-Bonsai-4B is the accuracy-per-byte winner above 1 GiB. 83.0% from only 1.1 GiB, within 2 points of Qwen3.5-4B at 40% of the weight size.

they show strong promise on edge devices and where disk space is limited. I think this lab is worth watching.

John Ternus to become Apple CEO

How to Make a Fast Dynamic Language Interpreter

Jujutsu megamerges for fun and profit

Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving

Kimi vendor verifier – verify accuracy of inference providers

Soul Player C64 – A real transformer running on a 1 MHz Commodore 64

Ternary Bonsai: Top Intelligence at 1.58 Bits

ggsql: A Grammar of Graphics for SQL

Quantum Computers Are Not a Threat to 128-Bit Symmetric Keys

OpenAI ad partner now selling ChatGPT ad placements based on “prompt relevance”

All phones sold in the EU to have replaceable batteries from 2027

Deezer says 44% of songs uploaded to its platform daily are AI-generated

Kefir C17/C23 Compiler

Modern Rendering Culling Techniques

Brussels launched an age checking app. Hackers took 2 minutes to break it

Even 'uncensored' models can't say what they want

Japan's Cherry Blossom Database, 1,200 Years Old, Has a New Keeper

Monero Community Crowdfunding System

Zero-Copy Pages in Rust: Or How I Learned to Stop Worrying and Love Lifetimes

WebUSB Extension for Firefox

M 7.4 earthquake – 100 km ENE of Miyako, Japan

Bloom (YC P26) Is Hiring

F-35 is built for the wrong war

Year of the IPv6 Overlay Network

10 years ago, someone wrote a test for Servo that included an expiry in 2026

Atlassian enables default data collection to train AI

Sauna effect on heart rate

Kimi K2.6: Advancing open-source coding

Writing string.h functions using string instructions in asm x86-64 (2025)

I learned Unity the wrong way

Ternary Bonsai: Top Intelligence at 1.58 Bits

Comments

John Ternus to become Apple CEO

How to Make a Fast Dynamic Language Interpreter

Jujutsu megamerges for fun and profit

Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving

Kimi vendor verifier – verify accuracy of inference providers

Soul Player C64 – A real transformer running on a 1 MHz Commodore 64

Ternary Bonsai: Top Intelligence at 1.58 Bits

ggsql: A Grammar of Graphics for SQL

Quantum Computers Are Not a Threat to 128-Bit Symmetric Keys

OpenAI ad partner now selling ChatGPT ad placements based on “prompt relevance”

All phones sold in the EU to have replaceable batteries from 2027

Deezer says 44% of songs uploaded to its platform daily are AI-generated

Kefir C17/C23 Compiler

Modern Rendering Culling Techniques

Brussels launched an age checking app. Hackers took 2 minutes to break it

Even 'uncensored' models can't say what they want

Japan's Cherry Blossom Database, 1,200 Years Old, Has a New Keeper

Monero Community Crowdfunding System

Zero-Copy Pages in Rust: Or How I Learned to Stop Worrying and Love Lifetimes

WebUSB Extension for Firefox

M 7.4 earthquake – 100 km ENE of Miyako, Japan

Bloom (YC P26) Is Hiring

F-35 is built for the wrong war

Year of the IPv6 Overlay Network

10 years ago, someone wrote a test for Servo that included an expiry in 2026

Atlassian enables default data collection to train AI

Sauna effect on heart rate

Kimi K2.6: Advancing open-source coding

Writing string.h functions using string instructions in asm x86-64 (2025)

I learned Unity the wrong way