frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Compression Is Prediction

https://ngrok.com/blog/compression-is-prediction
36•nikolay•51m ago

Comments

Muhammad523•35m ago
I was rushing to post this and then found out somebody had already
andai•17m ago
You snoze, you loze!

https://www.youtube.com/watch?v=B6u-FPskfAE

sheeeeesh•28m ago
Grant Sanderson has an excellent video on the same topic [0]. It's part of a series that is ongoing.

[0] https://youtu.be/l6DKRf-fAAM?si=yyLWq8x4sSRkWd98

zahrevsky•26m ago
I wonder if the author of the article knew about the series, or do they both just independently came across this topic to talk about it.
cyanydeez•7m ago
it was vaguely in my understanding of information & intelligence with compression; it was also brought up in several of the initial trials against AI companies where they discussed how the AI is akin to compression.

So they're both sourcing a bit broader zeitgeist.

throwaway_7274•27m ago
This perspective is a useful source of intuition against the “LLMs can’t have new ideas, they’re just next-token-predictors” style arguments. What if you shift your perspective to thinking of training as optimization over a vast parametrized family of compression algorithms? Well, it suddenly looks a lot more plausible that “new” “ideas” can emerge from that process!
throwaway_7274•25m ago
Incidentally, the relationship is bidirectional. You can try it out just for fun. zstd is a pretty crappy language model :)
glial•10m ago
> it suddenly looks a lot more plausible that “new” “ideas” can emerge from that process

This is not intuitive to me. It seems like a "new idea" is something that (almost by definition) isn't in the training set. Can you elaborate a bit?

Edit: but perhaps a good model could arise from training, which would be a good idea in the sense that parsimonious ideas are good scientific ideas.

deepsun•26m ago
> compressors and LLMs

Why only LLMs? All statistical models are compressor. You can say "model" and "compressor" are synonyms.

Article does not mention "embeddings" at all, even though it's commonly viewed as a compression method. Also "encoder" part on "auto-encoders".

andai•22m ago
See also: Bellard's Lossless Data Compression With Neural Networks

https://news.ycombinator.com/item?id=19589848

https://news.ycombinator.com/item?id=27244004

pjankiewicz•22m ago
I was thinking about the same topic and the conclusion can be wrong. LLMs are compressors, but compressors are not LLMs. Mixing this can let you believe that you can use a compressor to do the same thing as LLMs, which you cannot.

Specifically I was thinking about a way to inject knowledge into LLMs training by using statistical properties of text in such a way that you don't have to train the LLM to achieve some level of predictions. There are actually some papers that inject n-grams statistics as a part of the neural network weights.

Legend2440•10m ago
>Mixing this can let you believe that you can use a compressor to do the same thing as LLMs, which you cannot.

You can, actually! Any compressor can be losslessly converted into a generator, and vice versa.

Traditional compressors like gzip are of course very simple and can only replicate rough patterns from the input. But they are technically doing the same thing.

davmre•7m ago
Any compressor actually can be used, trivially, as an autoregressive language model.

Given a context (for LLMs, this would include the entire pretraining dataset, plus the prompt), you compress `context + next_token` for every possible next token. The tokens that co-compress best with the existing context are the 'least surprising' continuations. Choose one of them and iterate.

You can easily generate text with gzip this way. It won't be very good text, because gzip compression is not as sophisticated as a transformer + SGD, but the principle is the same.

aaroninsf•6m ago
That sounds like boostrapping the weights involved in early layers, to obviate the need for those layers to learn (optimize) for the distribution in the training set.

Makes me wonder idly, - is this conceptually akin in some sense to a "universal grammar," and if so - with a broad enough training set, is there a latent durable universal grammar that might be similarly recovered and injected to the benefit of all training, - does that grammar go beyond morphological/syntactical/grammatical features, into e.g. semantics and pragmatics

variadix•22m ago
This is a lot less surprising when you learn how non-LZ compressors work, that is, by modeling a probability distribution and using those probabilities to encode information in the minimum number of bits required to transmit the data. A less obvious conclusion is that LZ compressors do this to implicitly, the length of each symbol they could emit (literal or match, etc.) can be converted to the probability distribution the LZ compressor induces, since the number of bits to encode the symbol is related to its probability by the information content.
sethev•14m ago
This immediately reminded me of the Hutter Price (http://prize.hutter1.net/) - a contest that has run since 2005(?) based on the premise that compression is closely related to intelligence.
ssivark•13m ago
[delayed]
adamgordonbell•10m ago
Small world. I just did a podcast on this same topic, but coming at it from a different direction, ie. me and my neighbor trying to beat the hutter prize for compression.

Hutter Prize being where you are paid if you can compress wikipedia small enough. LLMs do very well at that, if, big if, you ignore the cost of initial weights.

https://corecursive.com/the-hutter-prize/

http://prize.hutter1.net/

https://github.com/hkust-nlp/llm-compression-intelligence

j-pb•7m ago
I always feel like something is missing when people say that compression and prediction are two sides of the same coin.

It feels more like a trinity: compression, prediction, indexing

In each case you try to recognise (re)usable structure.

Self-indexing succinct data-structures are a good example of the third side of the coin.

Show HN: Cut LLM turns in MCP interactions by 75%+

https://github.com/Tura-AI/tura
3•turaainet•2m ago•0 comments

.com now accounts for less than 48% of YC portfolio company domains

https://www.orangecrumbs.com/stories/yc-domains
2•oyster143•2m ago•0 comments

4,099 Qubits: The Myth and Reality of Breaking RSA-2048 with Quantum Computers

https://postquantum.com/post-quantum/4099-qubits-rsa/
2•Anon84•3m ago•0 comments

Boundaryguard – detect invisible Unicode/Trojan Source attacks in CI

https://github.com/000wq123/boundaryguard
2•davor1023•3m ago•0 comments

Lightweight, zero-build reactivity library build for MPA's

https://github.com/marsbos/flynt.js
2•harrald•3m ago•0 comments

Procedural Pixel-Art Grass in Rust and WebGPU

https://jarl-game.com/blog/2d-gpu-grass-rendering/
3•credo_dev•10m ago•0 comments

DFlash changes what tokens per second means

https://piszczek.pl/blog/dflash-changes-what-tokens-per-second-means
2•pich•12m ago•0 comments

Nix didn't hang it was evaluating the world

https://www.labcraft.dev/blog/nix-didnt-hang-it-was-evaluating-the-world
4•anandsuresh•14m ago•0 comments

CoreWeave edges past quarterly revenue estimates

https://www.reuters.com/technology/coreweave-edges-past-quarterly-revenue-estimates-2026-08-11/
2•ashurandi•16m ago•0 comments

The FastLanes Unified Transport Layout

https://blog.dave.tf/post/fastlanes-utl/
2•raggi•18m ago•0 comments

Google's Gemini app surges to one billion users

https://techcrunch.com/2026/08/11/googles-gemini-app-surges-to-one-billion-users/
3•speckx•19m ago•0 comments

Becoming physically immune to brute-force attacks (2021)

https://seirdy.one/posts/2021/01/12/password-strength/
2•downbad_•20m ago•0 comments

3D Interactive Map of the Silicon Valley

https://www.levels.fyi/atlas
2•zuhayeer•22m ago•1 comments

Bumble gives up on women making the first move

https://techcrunch.com/2026/08/11/bumble-ditches-its-rule-that-kept-men-from-making-the-first-move/
5•BlueBerry2001•24m ago•1 comments

uBlock Origin no longer supporting Facebook

https://reddit.com/r/uBlockOrigin/comments/1vgcjg5/about_disgusting_facebook_devs/
3•Cider9986•24m ago•1 comments

My agent hacked my gym

http://web.archive.org/web/20260516025532/https://www.affinda.com/expert-insights/when-my-ai-agen...
2•tmrtsmith•27m ago•1 comments

Previous-Token Prediction Based LLM Near-Exact Prompt Reconstruction

https://arxiv.org/abs/2607.29378
2•doener•27m ago•0 comments

Show HN: LidAwake – Keep your Mac awake with the lid closed (200 lines of Swift)

https://github.com/deezeddd/LidAwake-Mac
1•deezedd•28m ago•0 comments

Tiled Rasterization for Large DOM Captures

https://snapdom.dev/blog/huge-page-mosaic/
2•tinchox6•29m ago•0 comments

Introducing Grok Bot

https://xcancel.com/antibot/captcha
2•StanAngeloff•29m ago•0 comments

The Fanfare Around the Band Geese Was a Psyop

https://www.wired.com/story/geese-chaotic-good-marketing-industry-plant/
2•dbl000•29m ago•0 comments

Show HN: Personal AI Policy Directory

https://howiai.directory/
1•emilesilvis•30m ago•0 comments

Firefly Aerospace hits record $117.7M Q2 revenue, beating estimates by 34%

https://runtimewire.com/article/firefly-aerospace-hits-record-117-7m-q2-revenue-beating-estimates...
1•ryanmerket•30m ago•0 comments

Blend2D – High Performance 2D Vector Graphics

https://blend2d.com/
2•coffeeaddict1•31m ago•0 comments

Rocky Linux Founder Kurtzer Launches OpenWALDO to Open Up AI Training Data

https://techstrong.ai/articles/rocky-linux-founder-gregory-kurtzer-launches-openwaldo-to-open-up-...
3•CrankyBear•32m ago•0 comments

Glyph – Content intelligence without the baggage

https://github.com/Koda-OSS/Glyph
2•justinthemax•35m ago•0 comments

The brain may be about to have its Ozempic moment

https://economist.com/science-and-technology/2026/08/11/the-brain-may-be-about-to-have-its-ozempi...
32•andsoitis•40m ago•18 comments

Ask HN: How should I progress in my hacking journey?

1•anonomity•40m ago•0 comments

Trystero

https://github.com/dmotz/trystero
2•tibbar•41m ago•0 comments

Ancient Engineering Marvel Was Built Without Rulers

https://www.404media.co/no-bosses-ancient-engineering-marvel-was-built-without-rulers-study-sugge...
2•Jimmc414•42m ago•0 comments