frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

SOTA on the hardest AI memory benchmark (BEAM, 10M tokens), with a smaller model

2•johnnymakes•4h ago
Hey HN. I'm Johnny, founder of Exabase. We just hit the highest reported score on BEAM, the hardest AI memory benchmark, at every scale up to 10 million tokens. We also ran our evaluation using Gemini 3 Flash, when all previous leaders depended on a much larger model (Gemini 3 Pro).

At 10M tokens, the scale is vastly larger than any model's context window, so context stuffing isn't an option (aside from the fact that only about half of a large window can be effectively utilised without degradation). The only way to score well is recall that fundamentally works.

Our system (M-1) scored 76.9% at 100K, 75.0% at 1M, and 68.0% at 10M. Previous leaders were Hindsight (73.4%, 73.9%, 64.1%) and Honcho (63.0%, 63.1%, 40.6%), both using Gemini 3 Pro, while we used Flash.

We saw the competitive gap get wider at scale: 3.5 points ahead of Hindsight at 100K, 3.9 at 10M. The gap with Honcho goes from 13.9 to 27.4 points. As the corpus gets bigger, it filters out effective recall vs. brute-forcing / model capability.

M-1 also consumed about 20% fewer tokens per query than the next best system.

About the BEAM benchmark: BEAM tests ten memory abilities including some that other benchmarks don't cover: contradiction resolution, event ordering, and instruction following. The 10M token scale is vaguely equivalent to a year of long daily chats with an LLM.

Of course our system is still far from perfect, with strength in some categories (preference following, instruction following, summarization, abstention) all consistently above 90%, even at 10M token scale.

And our system shows weakness in others: for example, multi-session reasoning: 44.7% at 100K, collapsing to 9.6% at 10M. Although that challenge seems to be a general problem across memory systems at this scale, not M-1 specific. Something we'll continue to work on.

Methodology: We forked Hindsight's open-source benchmarking script, replaced the retrieval layer, and used the runner's prompt structure with minor adjustments for production use. Full methodology, results JSON for all three scales, and the prompt generator are linked in the paper (paper linked below).

Combined with our LongMemEval result (96.4%), M-1 is now the only system to hold SOTA across both major memory benchmarks at every scale, from 115K to 10M tokens.

Research paper: https://exabase.io/research/exabase-achieves-state-of-the-art-on-beam-benchmark

Happy to discuss architecture, the benchmark, scale challenges etc.

Comments

johnnymakes•4h ago
Research paper link: https://exabase.io/research/exabase-achieves-state-of-the-ar...
Chihuahua0633•3h ago
Imagine my disappointment to find out that this is _not_ about Satellites on the Air with a BEAM Antenna on the 10M frequency band ...
johnnymakes•1m ago
Haha – sorry to disappoint!

Ask HN: Crooked Timber showed showed me a virus captcha, What now?

39•Jgoauh•5h ago•44 comments

Tell HN: Our paid Claude AI subscription unavailable >1 week and no support

41•KellyCriterion•11h ago•21 comments

Why I prefer Opus 5 to Fable 5

19•novlrdotcom•9h ago•8 comments

Free AI visibility scanner for your website

2•hexalogic•3h ago•0 comments

Ask HN: How do you use Copilot in the CLI, App, and IDE?

3•EspressoGPT•6h ago•1 comments

SOTA on the hardest AI memory benchmark (BEAM, 10M tokens), with a smaller model

2•johnnymakes•4h ago•3 comments

Ask HN: How do you audit your app for compliance?

2•Luxter•4h ago•0 comments

Ask HN: How to rewrite `Claude.md` and install the skill for Opus5 and Fable5

5•hyhmrright•13h ago•6 comments

BMWs shows in-car ads for Spiderman

35•bigmattystyles•1d ago•17 comments

Can anti-fraud make large-scale attacks unprofitable?

3•jezzwar•8h ago•1 comments

Ask HN: Thoughts on AMD Ryzen AI MAX+ 395 for Local AI?

2•mempirate•9h ago•2 comments

Internet is no longer accessible?

43•cute_boi•17h ago•24 comments

Ask HN: Is there legal risk in AI memory?

2•sarjann•10h ago•0 comments

Ask HN: What do you use local models for?

3•tjomk•12h ago•2 comments

Ask HN: How to deal with security implications of running/installing projects?

12•johng•22h ago•11 comments

Goodbye POSIX, Hello Tea Room: Inside MinervaOS Architecture

10•xentnex•1d ago•6 comments

Ask HN: Why is every company encorporating AI everywhere?

24•vanessa1211•1d ago•39 comments

Ask HN: What are the most promising RL fields for a new master student?

34•zecice•2d ago•14 comments

Claude's code comments – too much or just enough?

8•lalaleslieeeee•16h ago•5 comments

Ask HN: What apps are you building?

14•totaldude87•17h ago•20 comments

Claude Code getting "API Error: 529 Overloaded"

5•croemer•1d ago•2 comments

Messaging platform I can self-host for my agents to communicate

2•giribisthebest•21h ago•3 comments

You've reached the end!