frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

SOTA on the hardest AI memory benchmark (BEAM, 10M tokens), with a smaller model

1•johnnymakes•43m ago
Hey HN. I'm Johnny, founder of Exabase. We just hit the highest reported score on BEAM, the hardest AI memory benchmark, at every scale up to 10 million tokens. We also ran our evaluation using Gemini 3 Flash, when all previous leaders depended on a much larger model (Gemini 3 Pro).

At 10M tokens, the scale is vastly larger than any model's context window, so context stuffing isn't an option (aside from the fact that only about half of a large window can be effectively utilised without degradation). The only way to score well is recall that fundamentally works.

Our system (M-1) scored 76.9% at 100K, 75.0% at 1M, and 68.0% at 10M. Previous leaders were Hindsight (73.4%, 73.9%, 64.1%) and Honcho (63.0%, 63.1%, 40.6%), both using Gemini 3 Pro, while we used Flash.

We saw the competitive gap get wider at scale: 3.5 points ahead of Hindsight at 100K, 3.9 at 10M. The gap with Honcho goes from 13.9 to 27.4 points. As the corpus gets bigger, it filters out effective recall vs. brute-forcing / model capability.

M-1 also consumed about 20% fewer tokens per query than the next best system.

About the BEAM benchmark: BEAM tests ten memory abilities including some that other benchmarks don't cover: contradiction resolution, event ordering, and instruction following. The 10M token scale is vaguely equivalent to a year of long daily chats with an LLM.

Of course our system is still far from perfect, with strength in some categories (preference following, instruction following, summarization, abstention) all consistently above 90%, even at 10M token scale.

And our system shows weakness in others: for example, multi-session reasoning: 44.7% at 100K, collapsing to 9.6% at 10M. Although that challenge seems to be a general problem across memory systems at this scale, not M-1 specific. Something we'll continue to work on.

Methodology: We forked Hindsight's open-source benchmarking script, replaced the retrieval layer, and used the runner's prompt structure with minor adjustments for production use. Full methodology, results JSON for all three scales, and the prompt generator are linked in the paper (paper linked below).

Combined with our LongMemEval result (96.4%), M-1 is now the only system to hold SOTA across both major memory benchmarks at every scale, from 115K to 10M tokens.

Research paper: https://exabase.io/research/exabase-achieves-state-of-the-art-on-beam-benchmark

Happy to discuss architecture, the benchmark, scale challenges etc.

Comments

johnnymakes•41m ago
Research paper link: https://exabase.io/research/exabase-achieves-state-of-the-ar...

Xenharmlib (music theory library) adds support for Just Intonation

https://xenharmlib.readthedocs.io/en/latest/whats_new_0_4_0.html
1•retooth•26s ago•0 comments

Destructive Convenience

https://volpeon.ink/notebook/dangerous-convenience/
1•mintplant•57s ago•0 comments

Spam - a System PAckage Manager for folks who use multiple platforms

https://codeberg.org/aol/spam
1•mixmastamyk•1m ago•0 comments

"Uncensored" open LLMs are measurably more optimistic than their base models

https://arxiv.org/abs/2607.17427
1•oleczek•1m ago•0 comments

WebKit Features for Safari 26.6

https://webkit.org/blog/18178/webkit-features-for-safari-26-6/
1•ksec•2m ago•0 comments

The Sticky Mark-Bit Algorithm for GCs

https://wingolog.org/archives/2022/10/22/the-sticky-mark-bit-algorithm
1•achierius•2m ago•0 comments

Home Office used 'AI hallucinated' information to refuse asylum claim, judge

https://www.theguardian.com/uk-news/2026/jul/28/home-office-used-ai-hallucinated-information-to-r...
2•sbulaev•3m ago•0 comments

Show HN: A search engine for one-word .com domains

https://www.muffin-company.com/word-domains
2•garyhbutton•5m ago•0 comments

Please, Don't Screw Up Neuromancer

https://freddiedeboer.substack.com/p/please-dont-screw-up-neuromancer
4•paulpauper•6m ago•0 comments

Debating Housing and More with Steve Hilton

https://www.richardhanania.com/p/debating-housing-and-more-with-steve
3•paulpauper•6m ago•0 comments

I showed an AI an image it couldn't see – then caught it lying about what it saw

https://paulshepherd.com/conversations/session-1.html
3•LifeWithGlee•7m ago•1 comments

Paged Out – Issue #9

https://pagedout.institute/webview.php?issue=9&page=1
3•birdculture•7m ago•0 comments

Claude Opus 5: Model Welfare

https://thezvi.substack.com/p/claude-opus-5-model-welfare
4•paulpauper•7m ago•0 comments

Model Context Protocol Grows Up

https://techstrong.ai/articles/model-context-protocol-grows-up/
3•CrankyBear•8m ago•0 comments

Built Once, Serves Many: How We Rebuilt Experiment Analysis Platform at Roku

https://engineering.roku.com/built-once-serves-many-how-we-rebuilt-experiment-analysis-platform-a...
3•milindnandankar•8m ago•1 comments

Removing Apache Log Noise

https://j11g.com/removing-apache-log-noise
2•speckx•8m ago•0 comments

You Could Have Come Up with Kimi Delta Attention

https://blog.doubleword.ai/you-could-have-come-up-with-kimi-delta-attention
2•AnhTho_FR•8m ago•0 comments

Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents

https://arxiv.org/abs/2607.24625
2•zhinit•10m ago•0 comments

Pay Gruber

https://ninjasandrobots.com/pay-gruber
3•nate•12m ago•0 comments

Mysterious Human Relative Was Bigger Than Previously Thought

https://nautil.us/mysterious-human-relative-was-bigger-than-previously-thought-1283071
2•Brajeshwar•14m ago•0 comments

How Do I Profile eBPF Code?

https://naveensrinivasan.com/posts/2026-07-22-how-do-i-profile-ebpf-code/
9•snaveen•15m ago•0 comments

Show HN: XY – Fast, composable, GPU-accelerated charts, written in Rust

https://github.com/reflex-dev/xy
13•apetuskey•16m ago•1 comments

'It sounds like someone set up a vacuum ': Michigan residents sue AI DC

https://www.tomshardware.com/tech-industry/data-centers/it-sounds-like-someone-set-up-a-vacuum-li...
3•LaSombra•17m ago•0 comments

The making of Don Matrelli's Legacy, a mod for Grand Prix Circuit (part I)

https://marnetto.net/2026/07/18/dml-making-of-1
2•speckx•17m ago•0 comments

Code for Real: Teach kids to program with drag and drop blocks

https://code-for-real.pages.dev/get
2•eappleby•17m ago•2 comments

Efficient code for inverting the (log) gamma function

https://www.johndcook.com/blog/2026/07/28/inverse-factorial-improved/
2•ibobev•18m ago•0 comments

Draft: Complete redesign ( 349) · GNOME / Document Scanner

https://gitlab.gnome.org/GNOME/simple-scan/-/merge_requests/349
2•LaSombra•18m ago•0 comments

A Crosswording Miscellany

https://quuxplusone.github.io/blog/2026/07/28/puzzle-miscellany/
2•ibobev•19m ago•0 comments

The inliner is yielding benefits for ZJIT

https://railsatscale.com/2026-07-28-the-inliner-is-yielding-benefits/
2•ibobev•19m ago•0 comments

Delayed Gratification – Proud to Be 'Last to Breaking News'

https://www.slow-journalism.com/
5•speerer•20m ago•0 comments