60% Fable cost cut by converting code to images and having the model OCR it

44•dimitropoulos•2h ago

Comments

dimitropoulos•2h ago

there's also a DeepSeek whitepaper on this technique https://www.seangoedecke.com/text-tokens-as-image-tokens

genxy•1h ago

This seems like a pricing hack that burns resources, that when the loophole gets closed the price of OCR will have to rise?

ricardobeat•1h ago

It’s not a loophole, it just happens that encoding information as optical tokens is much more efficient than text.

guardiangod•40m ago

Truly a picture is worth a thousand words.

TZubiri•32m ago

Of course it isn't

A text encoding uses 8bits per character on average, tokenization further compresses that

An image font would be 25 bits if 5x5, and most fonts are 12 pixels high

Of course it isn't efficient, this is a pricing inefficiency and a hack to exploit it (even the author describes it as an exploit)

legel•10m ago

You are wrong.

Text tokens are high-dimensional vectors, not 8 bits per character. Every token has a deep embedding, e.g. 1024 float values per text token.

DeepSeek-OCR proved 10x+ compression from visual embedding of text, which was a groundbreaking result. [1]

Very cool to see OP's project hacking on this principle. It's still not lossless, as noted in the github, but is a promising research direction.

[1] https://github.com/deepseek-ai/DeepSeek-OCR/blob/main/DeepSe...

geor9e•6m ago

If it's not intuitive by 8 bit characters compress better than 8x8 pixel squares, then step back and think about it another way - ask which scenario is more likely:

Some random person discovered a 60% across the board gain in all LLMs, using an extremely simple trick that none of the labs noticed in all these years of multi-trillion dollar growth

Anthropic's marketing team might not have priced images on par with text in their rush to drive growth via money losing offerings

samrus•48m ago

Not really. They arent actually using more resources this way either. This might be a fundamental inefficiency thats being removed

It kinda makes sense too. Because while people do read code word by word, we often "glance over" it and do roughly pattern recognition on it to know what it does. Only homing in on something when we need to answer a specific question. I think humans kinda naturally do this exploit anyway

aabhay•1h ago

Ahhh my eyes the vibe coded readme

mpalmer•34m ago

What, you don't like your caveats to be honest?

lpellis•48m ago

I tried the same thing last year (with openai models), back then it worked to reduce prompt tokens, but you needed way more completion tokens, ultimately more expensive (and slower) https://pagewatch.ai/blog/post/llm-text-as-image-tokens/

aabhay•34m ago

In Gemini at least, if you look at how they process PDFs, they do an OCR and then feed the text + image to the model, without charging you for the text tokens (I believe).

So my guess is that Claude’s backend is doing the same — so this hack is probably more of a loophole in token accounting that might get closed if Claude is doing what Gemini does

dippogriff•33m ago

I want to see more text-free foundation models

puppycodes•26m ago

That is hilarious and an amazing find.

__hugues•7m ago

seems really dumb and like it would need to violate basic information theory to work?

input tokens are cheaper than output tokens. seems like it would maybe reduce input tokens at the expense of many more output tokens if you're actually triggering OCR via thinking?

himata4113•5m ago

Claude, please stop trying to memorize random crap

The Life and Times of Maxis, Part 1: SimEverything

Half-Baked Product

Jamesob's guide to running SOTA LLMs locally

International chess federation sanctions Kramnik

Factories Are Just Rooms

Hunting a 16-year-old SQLite WAL bug with TLA+

PostgreSQL and the OOM Killer: Why We Use Strict Memory Overcommit

My Dad Helped Build North America's Oat Supply Chain: Can It Be Remade?

Valve open source the Steam Machine e-ink screen so you can make your own

The Fall and Rise of Screwworm

Best Simple System for Now

Wordgard: The new in-browser rich-text editor from the creator of ProseMirror

America, 1926: What a Forgotten 100-Year-Old Report Says About Who We Are

Right to Local Intelligence

Supersonic flight returning to US after half-century ban

CarPlay Is Additive

Give Smart People the Tools to Do Smart Things

Anatomy of Persistent Memory's 3 Layers: Comparing ContextNest, Mem0 and Zep

Show HN: Mcpsnoop – Wireshark for MCP (transparent proxy and live TUI)

60% Fable cost cut by converting code to images and having the model OCR it

US residents angry datacenters 'shoved down our throats' are recalling officials

The Safari MCP server for web developers

How working with a blind client revealed invisible accessibility gaps

crustc: entirety of `rustc`, translated to C

Commodore 64 Basic for PostgreSQL

Markets are competitive if and only if P != NP

Reality has a surprising amount of detail (2017)

Quake in 13 Kilobytes (2021)

Program-as-Weights: A Programming Paradigm for Fuzzy Functions

60% Fable cost cut by converting code to images and having the model OCR it

Comments

Claude, please stop trying to memorize random crap

The Life and Times of Maxis, Part 1: SimEverything

Half-Baked Product

Jamesob's guide to running SOTA LLMs locally

International chess federation sanctions Kramnik

Factories Are Just Rooms

Hunting a 16-year-old SQLite WAL bug with TLA+

PostgreSQL and the OOM Killer: Why We Use Strict Memory Overcommit

My Dad Helped Build North America's Oat Supply Chain: Can It Be Remade?

Valve open source the Steam Machine e-ink screen so you can make your own

The Fall and Rise of Screwworm

Best Simple System for Now

Wordgard: The new in-browser rich-text editor from the creator of ProseMirror

America, 1926: What a Forgotten 100-Year-Old Report Says About Who We Are

Right to Local Intelligence

Supersonic flight returning to US after half-century ban

CarPlay Is Additive

Give Smart People the Tools to Do Smart Things

Anatomy of Persistent Memory's 3 Layers: Comparing ContextNest, Mem0 and Zep

Show HN: Mcpsnoop – Wireshark for MCP (transparent proxy and live TUI)

60% Fable cost cut by converting code to images and having the model OCR it

US residents angry datacenters 'shoved down our throats' are recalling officials

The Safari MCP server for web developers

How working with a blind client revealed invisible accessibility gaps

crustc: entirety of `rustc`, translated to C

Commodore 64 Basic for PostgreSQL

Markets are competitive if and only if P != NP

Reality has a surprising amount of detail (2017)

Quake in 13 Kilobytes (2021)

Program-as-Weights: A Programming Paradigm for Fuzzy Functions