frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

"A milion token context" Big AI says. But the model is accurate for 2-4K tokens

https://unagent.eu/2025/04/22/misleading-promises-of-long-context-llm/
2•kzawpl•1y ago

Comments

kzawpl•1y ago
Over last two years there were claims of better long context capabilities for LLM, but that is often tested on exact text search. New benchmark called NoLiMa shows that long context capability of LLM is still poor, if you want LLM to perform some abstraction and reasoning.
vessenes•1y ago
Meh. NoLima is helpful, in that it shows what we all "feel" working with models -- there's a marked dropoff in accuracy and intelligence as we get past 4-32k of context, depending on the model.

But, it seems unreasonable to be super worried about this -- a year or two ago, models couldn't easily find needles in haystacks of long context. As training and test strategies delivered trainable content, this became a thing that could be done perfectly across millions of tokens of context. There has not been a good way to incentivize models to do anything more but remember locations yet.

We are (mostly) paying the full costs of attending to the entire context in current architectures, and it seems pretty reasonable that we will therefore be able to train those architectures to more fully attend across context if we get the right training data into (ideally) an RL loop.

NoLima is an okay test, but I think the most recent OpenAI tests are significantly better and quite interesting; OpenAI-MRCR and Graphwalks are both super smart ideas about how to programmatically generate data that is easy to evaluate and forces better cross context attention.

From their 4.1 announcement: Graphwalks fills the context window with a directed graph composed of hexadecimal hashes, and then asks the model to perform a breadth-first search (BFS) starting from a random node in the graph. We then ask it to return all nodes at a certain depth.

MRCR asks for direct quotes at semantically identified locations in the text, e.g. poems about tapirs, bears and ballerinas, as well as stories about tapirs, bears and ballerinas are generated, perhaps fifty each. The system is asked "give me the third poem about tapirs". This requires counting, conceptual attention, and also distinguishing between stories and poems.

They only test their own models on MRCR for the benchmark graph, but it's still worth reviewing: the accuracy curves are super interesting. https://openai.com/index/gpt-4-1/

Show HN: Hntui – A TUI for Hacker News

https://github.com/ahmd-sh/hntui
1•ahmd-sh•25s ago•0 comments

The End of Hand-Written Code: A Conversation with David Heinemeier Hansson

https://thoughteconomics.com/david-heinemeier-hansson/
1•cebert•2m ago•0 comments

Neural Drive

https://mlx-optiq.com/blog/neural-drive-game
1•codelion•3m ago•0 comments

Chatbots give a narrow slice of knowledge:Researchers warn of knowledge collapse

https://news.ku.dk/all_news/2026/09/ai-chatbots-give-us-a-narrow-slice-of-knowledge-researchers-w...
1•DeepLogin•4m ago•0 comments

China Broadens Travel Curbs to Encompass Family of Top AI Talent

https://www.bloomberg.com/news/articles/2026-09-28/china-broadens-travel-curbs-to-encompass-famil...
1•sbulaev•4m ago•0 comments

Cohesix – find out what happened to a local AI job

https://github.com/lukeb-aidev/cohesix/releases/tag/v1.2.0
1•Cohesix•6m ago•0 comments

Richardson Maturity Model, Steps Toward the Glory of REST

https://martinfowler.com/articles/richardsonMaturityModel.html
1•vladde•9m ago•1 comments

Nvidia Launches Open Agent Safety Platform

https://mrkt30.com/nvidia-open-agent-safety-platform-openshell-sentry/
1•newscomAI•9m ago•0 comments

Vite+ 1.0 Is Out

https://voidzero.dev/posts/announcing-vite-plus-1-0
1•vanyle•9m ago•0 comments

The Largest Roman Mosaic Ever Excavated Opens to the Public

https://www.thisiscolossal.com/2026/09/largest-roman-mosaic-museum-rome/
1•surprisetalk•11m ago•0 comments

Space X Startship first orbital flight

https://www.nytimes.com/2026/09/28/science/space/spacex-starship-first-orbital-flight.html
1•ltononro•11m ago•0 comments

SlopTotal, a Self-hosted AI text detector that runs 23 open models

https://github.com/pablocaeg/sloptotal
3•sloptotal•11m ago•0 comments

Geely to Buy 30% Stake in Rival Nio's Battery-Swapping Unit

https://www.bloomberg.com/news/articles/2026-09-28/geely-to-buy-30-stake-in-rival-nio-s-battery-s...
1•teleforce•12m ago•0 comments

Fingerprints of Jev

https://jdhornsby.com/fingerprints-of-jev/
1•time0ut•13m ago•0 comments

Why Software Factories Fail [video]

https://www.youtube.com/watch?v=Ib5GBkD555M
1•mpweiher•14m ago•0 comments

Intro to Clojure Workshop [video]

https://www.youtube.com/watch?v=KwM1c7vb-fg
1•doubleg•14m ago•0 comments

Is an Agentic Bank Run Coming?

https://www.apollo.com/wealth/insights-news/insights/daily-spark/is-an-agentic-bank-run-coming
2•theanonymousone•15m ago•0 comments

My coding agent now records the demo video for its own PRs

https://github.com/half144/cutaway
1•half144•15m ago•0 comments

#59 Hans Kelsen and Carl Schmitt: Friend and Enemy

https://voelkerrechtsblog.org/59-hans-kelsen-and-carl-schmitt-friend-and-enemy/
1•jruohonen•17m ago•0 comments

Hospitals use AI to find more things to bill for. Insurers use AI to deny them.

https://twitter.com/HedgieMarkets/status/2104266774766039301
2•MrBuddyCasino•18m ago•0 comments

Parisians on Hunt for Baguettes as Bakers Get Nod to Take Vacation (2015)

https://www.npr.org/sections/thesalt/2015/08/25/433005099/parisians-on-hunt-for-baguettes-as-bake...
1•thunderbong•20m ago•0 comments

How NOT to detect residential proxies

https://blog.truesign.ai/posts/how-NOT-to-detect-residential-proxies/
6•kanokano•20m ago•2 comments

Oracle triggers 'force majeure' on data center project over power delays

https://www.reuters.com/business/oracle-cites-force-majeure-shield-itself-controversial-data-cent...
1•andyjohnson0•20m ago•0 comments

The case for Wile E. Coyote, engineer: how it shaped the way we build

https://school.coyotiv.com/blog-wile-e-coyote-engineer-en
1•arbayi•21m ago•0 comments

Show HN: Skill that lets you run other Skills as Jinja templates

https://github.com/electronick1/LLAJinja
2•olekrk•21m ago•0 comments

Corporate America embraces cheaper 'open' AI models

https://www.ft.com/content/d9de4776-1fc9-4f2b-aaaf-9961c35d8acd
1•ostenbom•22m ago•0 comments

A Replacement for Bert: Introducing ModernBERT

https://www.answer.ai/posts/2024-12-19-modernbert.html
1•6bitquant•23m ago•0 comments

Vanguard: Anti-Boost, Pro-Skill

https://www.riotgames.com/en/news/vanguard-anti-boost
1•kabakabakaba•24m ago•1 comments

Applying Deming's Continuous Improvement to Multi-Agent AI Systems

https://medium.com/@zrkjsy/deming-in-the-machine-running-an-ai-assisted-film-production-on-contin...
1•taivare•24m ago•0 comments

Show HN: Quick Decision - Select text, get a yes/no verdict (Jev)

https://www.getopenclip.app/extensions/com.openclip.quickdecision
1•ganeshmshetty•24m ago•0 comments