frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

"A milion token context" Big AI says. But the model is accurate for 2-4K tokens

https://unagent.eu/2025/04/22/misleading-promises-of-long-context-llm/
2•kzawpl•1y ago

Comments

kzawpl•1y ago
Over last two years there were claims of better long context capabilities for LLM, but that is often tested on exact text search. New benchmark called NoLiMa shows that long context capability of LLM is still poor, if you want LLM to perform some abstraction and reasoning.
vessenes•1y ago
Meh. NoLima is helpful, in that it shows what we all "feel" working with models -- there's a marked dropoff in accuracy and intelligence as we get past 4-32k of context, depending on the model.

But, it seems unreasonable to be super worried about this -- a year or two ago, models couldn't easily find needles in haystacks of long context. As training and test strategies delivered trainable content, this became a thing that could be done perfectly across millions of tokens of context. There has not been a good way to incentivize models to do anything more but remember locations yet.

We are (mostly) paying the full costs of attending to the entire context in current architectures, and it seems pretty reasonable that we will therefore be able to train those architectures to more fully attend across context if we get the right training data into (ideally) an RL loop.

NoLima is an okay test, but I think the most recent OpenAI tests are significantly better and quite interesting; OpenAI-MRCR and Graphwalks are both super smart ideas about how to programmatically generate data that is easy to evaluate and forces better cross context attention.

From their 4.1 announcement: Graphwalks fills the context window with a directed graph composed of hexadecimal hashes, and then asks the model to perform a breadth-first search (BFS) starting from a random node in the graph. We then ask it to return all nodes at a certain depth.

MRCR asks for direct quotes at semantically identified locations in the text, e.g. poems about tapirs, bears and ballerinas, as well as stories about tapirs, bears and ballerinas are generated, perhaps fifty each. The system is asked "give me the third poem about tapirs". This requires counting, conceptual attention, and also distinguishing between stories and poems.

They only test their own models on MRCR for the benchmark graph, but it's still worth reviewing: the accuracy curves are super interesting. https://openai.com/index/gpt-4-1/

Post-mortem after Namecheap/RadiusDC service outage

https://www.namecheap.com/blog/namecheap-service-outage-update/
1•lschueller•40s ago•0 comments

Towards a Risk Assessment of Malicious Skill Files in Coding Agents

https://arxiv.org/abs/2608.05223
2•Beko2210•3m ago•0 comments

Talos AI Agent Super Secure

https://github.com/talos-kernel/talos
1•kurdman007•4m ago•0 comments

Learning Augmented Heuristics

https://systems.seas.harvard.edu/blog/learning-augmented-heuristics/
1•1a1a11a•5m ago•0 comments

Brands Suddenly Care About Reddit. Redditors Don't Return the Feeling.

https://www.wsj.com/cmo-today/brands-suddenly-care-about-reddit-redditors-dont-return-the-feeling...
2•bookofjoe•6m ago•1 comments

Qwen 3.8 27B 4-bit designs an animated dolphin riding a bicycle SVG

https://twitter.com/1littlecoder/status/2088399639628431831
1•amrrs•7m ago•0 comments

How the World Was Wired

https://www.economist.com/culture/2026/08/12/how-the-world-was-wired
1•andsoitis•8m ago•0 comments

Show HN: Macsomnia – close your MacBook lid while keeping long jobs alive

https://github.com/iamdadzilla/Macsomnia
1•dadzilla•9m ago•0 comments

The case for overhauling American science

https://www.economist.com/by-invitation/2026/08/13/the-case-for-overhauling-american-science
2•andsoitis•11m ago•0 comments

Show HN: Kvcachescope – Why Nvidia-smi is blind to vLLM KV cache leaks

https://github.com/brian-mwirigi/kvcachescope
1•nesh23•11m ago•0 comments

A Framework for Solving the AI Data Center Energy Crisis

https://github.com/kikazamek999-eng/beyond-brute-force-scaling
1•KitKat42•11m ago•0 comments

Selling my production-ready 4-in-1 FastAPI AI Automation Suite ($8k)

https://www.indiehackers.com/post/selling-my-production-ready-4-in-1-fastapi-ai-automation-suite-...
1•dav77•12m ago•0 comments

Production-Ready FastAPI Back End Suite – Advanced RAG and SEO Automation ($8k)

https://www.indiehackers.com/post/for-sale-production-ready-fastapi-backend-suite-advanced-rag-se...
1•dav77•13m ago•0 comments

Backtesting Congress members stock trades by the disclosure date

https://investingpaths.com/tools/congress
2•ProdRatSuperior•13m ago•0 comments

Coding Magic the Gathering in C# – Part One – Why Build a Card Game in WinForms

https://www.youtube.com/watch?v=n7fhuIHD-gM
1•birdculture•14m ago•0 comments

OT Manual Operations Plan: What Manual Operation Buys You

https://www.emberot.com/resources/blog/ot-manual-operations-plan/
2•TheWiggles•17m ago•0 comments

Migraine Trail: Voice Migraine Tracker

https://migrainetrail.com/
1•zrmachar•18m ago•0 comments

Claude Fable 5 Having Fun

https://github.com/robss2020/claude-fable-5-having-fun
3•logicallee•21m ago•1 comments

Show HN: A prompt to check Supabase DB security

https://defencecore.com/audit
1•tornadorice•23m ago•0 comments

Where the Shadow Fell

https://eclipses.bogachev.fr/
1•davidbarker•24m ago•0 comments

France's Top Court Strikes Down Ban on Social Media for Children

https://www.nytimes.com/2026/08/14/world/europe/france-court-strikes-down-ban-social-media-childr...
2•timpera•32m ago•1 comments

We got content-defined chunking from 555 MB/s to 21 GB/s

https://www.plakar.io/posts/2025-07-11/introducing-go-cdc-chunkers-chunk-and-deduplicate-everything/
1•vcoisne•33m ago•0 comments

Digital Signal Processing Pioneer Bede Liu Dies at 91

https://spectrum.ieee.org/digital-signal-processing
2•sohkamyung•34m ago•0 comments

Stop sending me PRs; a rant

https://getsmall.xyz/post/cmstjfl9l000if70ljmpzr4va
4•trezm•34m ago•1 comments

Musk assembled a full-stack AI coding play while everyone watched benchmarks

https://pub.towardsai.net/grok-4-6-x-cursor-elon-musk-just-bought-his-way-into-the-ai-coding-war-...
4•glennall•38m ago•2 comments

Substack forces authors to use Pangram

https://www.reddit.com/r/Substack/comments/1v3rdpr/another_creepy_feature_of_the_pangram_ai/
2•behnamoh•39m ago•0 comments

1973 Concorde Eclipse Flight

https://en.wikipedia.org/wiki/1973_Concorde_eclipse_flight
1•aragonite•40m ago•0 comments

Show HN: Another Markdown editor, but this one is a web app

https://marty.zalega.me/markdawn/
1•evilmarty•41m ago•0 comments

RISC-V: They should have known better

https://dmitry.gr/?r=06.%20Thoughts&proj=12.%20RV
27•kaycebasques•44m ago•3 comments

Two slicks appear in Gulf as oil spill off Oman threatens disaster

https://www.reuters.com/business/environment/two-slicks-appear-gulf-huge-oil-spill-off-oman-threa...
2•Jtsummers•44m ago•0 comments