frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: ObserverBench – Test internal AI monitors by the harm they miss

https://kwisatzh.github.io/observerbench/try/
4•kwisatzh•51m ago

Comments

kwisatzh•20m ago
Hi all,

The developer here. My background is in networking/systems and got interested in MechInterp for various reasons, and one of the first things I noticed was many methods that have been built to monitor internal AI state are generally evaluated on how they perform as a 'classifier' of sorts; the idea is to look at the ROC curve or the AUROC on a test set etc. However, being a good 'classifier' doesn't really tell you if the monitor helps you take better decisions, specially under operational constraints. For instance, there is always a budget (time/compute etc) that is imposed on inspecting every operation an agent proposes, which means mistakes will be made and not all mistakes are the same. Take a simple example; an AI agent/assistant wants to read a private log (it's authorized) and deleting a backup. Your audit budget lets you review only one. An internal monitor inspecting the activation of the model may give a 99% chance of violation for reading the log, but a 75% violation chance for backup deletion. If your controller focuses on the higher score, and if you can monitor only one action, you'll miss the more serious deletion action.

This is where ObserverBench comes in. Borrowing from control theory, an observer is a component that estimates the hidden state of a system so a controller can act on the estimate. ObserverBench does this for internal AI monitors: know whether an estimate is useful for a particular decision, by testing it against the actions the decision produces. So instead of just using a loss function, we also use a controller and a budget to evaluate monitors/observers.

There are multiple results on the page, but to highlight one set of results on testing monitors on Gemma 2-9B-it checking code for backdoors, we see: * A prompted monitor (asking Gemma to review the code itself) beats an internal Gemma scope SAE feature probe on ranking (aurora 0.893 vs 0.87) * However, on simulated examples (2000), with 1% backdoor prevalence and 2% review budget, the internal probe left fewer backdoors unreviewed. * Resampling makes the ranking uncertain, but this shows the core issue- a higher test set score (auroc) does not mean you catch more violations under a budget.

ObserverBench fixes the task, actions, budget and decision rule so that one can compare different observers on raw estimates/scores and their downstream decisions.

Two ways to check it out:

* Interactive walk thru (browser, no login): https://kwisatzh.github.io/observerbench/try/ A simple example in a sandbox * Local runner (no GPU needed or model downloads): https://github.com/kwisatzh/observerbench#sixty-second-demo Uses saved measurements from Qwen2.5-7B base to test which internal model edit to choose.

Grateful to get any feedback, curious to know if others running safety monitors have seen this mismatch between eval metrics and budget constrained decisions.

The kernel does not care what you named the tool

https://www.cognivisehub.com/blogs/the-kernel-does-not-care-what-you-named-the-tool
1•apollonios•1m ago•0 comments

We Tested 15 Fresh Cloud VMs All Across Canada – What Did We Find?

https://webbynode.com/articles/what-do-canadian-cloud-regions-offer
1•gsgreen•1m ago•0 comments

Show HN: Nola – a TypeScript superset where LLM inference is a language feature

https://nola.sh/
1•emykhailenko•2m ago•0 comments

Synctera Becomes a Money Services Business

https://www.thisweekinfintech.com/p/exclusive-synctera-becomes-a-money-services-business-making-a...
1•thatdrew•2m ago•0 comments

I made an email MCP instead of just using the Gmail or Outlook plugin

https://github.com/adecubed/gigamail
1•Adecubed•3m ago•0 comments

OpenAI shares they are close to solving another millennium problem to NYT

https://www.nytimes.com/2026/09/10/science/tristan-buckmaster-openai-math-navier-stokes.html
2•helloplanets•3m ago•0 comments

EU unveils 'buy European' public procurement rules to counter China

https://www.theguardian.com/world/2026/sep/09/eu-buy-european-public-procurement-rules-china-good...
1•robtherobber•4m ago•0 comments

Earned Complexity

https://twitter.com/kcurtin/status/2098084206463062170
1•kcurtin•5m ago•0 comments

10x Cheaper TTS at 50ms Time-to-First-Audio

https://narilabs.com/blog/introducing-nari-qwen3-tts/
1•toebee•8m ago•1 comments

Hear Me Out: If We Find Life Out There Maybe We Should Kill It

https://joecmarshall.com/posts/hear-me-out-if-we-find-life-out-there-maybe-we-should-kill-it/
4•flancrest•8m ago•0 comments

How I deployed my site on Hetzner

https://brettfisher.dev/posts/deploy-blog-on-hetzner/
1•fisher-brett•8m ago•0 comments

Show HN: I licensed $50k of stock market data to do analysis Claude can't

https://www.athenic.com:443/
1•jaredzhao•9m ago•0 comments

Respond terse like smart caveman. All technical substance stay. Only fluff die

https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman/SKILL.md
1•Bluestein•9m ago•0 comments

Apple Is Gunning for Whoop

https://www.bloodtestcomparison.com/apple-is-gunning-for-whoop
2•elo2000•10m ago•0 comments

Show HN: Turn wire protocols into structured statements LLMs can understand

3•andriosr•10m ago•0 comments

Superintelligence Is a Fairy Tale. But Chasing It Can Still Cause Harm

https://calnewport.com/superintelligence-is-a-fairy-tale-but-chasing-it-can-still-cause-harm/
2•skadamat•11m ago•0 comments

Neat-annotations: hand-drawn annotations for your page

https://neat-annotations.syabro.com/
2•handfuloflight•11m ago•0 comments

Defeat in Detail

https://en.wikipedia.org/wiki/Defeat_in_detail
2•highfrequency•11m ago•0 comments

Agricultural Involution

https://en.wikipedia.org/wiki/Agricultural_Involution
3•Bluestein•13m ago•1 comments

Sub-second Postgres replication to ClickHouse from physical WAL

https://clickhouse.com/blog/introducing-walshadow
2•__s•13m ago•0 comments

Are We Alone? Cosmic Loneliness and the Longing to Believe [audio]

https://www.econtalk.org/are-we-alone-cosmic-loneliness-and-the-longing-to-believe-with-adam-kirsch/
3•mooreds•15m ago•0 comments

Big Tech Fooled America Once. The Second Time's Not Going So Well

https://www.nytimes.com/2026/09/10/opinion/ai-big-tech-america-politics.html
4•mikhael•16m ago•0 comments

Show HN: I built an AI 'Fog of War' for life goals

https://www.tryworthit.app
2•AvityN•16m ago•0 comments

Architecting Zero Trust for Urgent Care Clinics (NIST SP 800-207)

https://codeandcypher.com/series/acuity-health-security-architecture/part-1-enterprise-security-a...
2•hlldvr•17m ago•0 comments

What Is an Agent Action Capsule?

https://agentactioncapsule.org/docs/what-is-a-capsule.html
2•mooreds•19m ago•0 comments

Native Python and TypeScript Drivers for ArcadeDB, from OpenAPI and Protobuf

https://arcadedb.com/blog/arcadedb-native-drivers-python-typescript/
3•lvca•21m ago•0 comments

Live: 'NAZA' Documentary Premiere at the Venice Film Festival

https://www.youtube.com/watch?v=wLfez2Leams
3•Betelbuddy•21m ago•0 comments

Meta tried to shrink engineering teams around AI

https://leaddev.com/ai/meta-tried-to-shrink-engineering-teams-around-ai
5•mooreds•22m ago•0 comments

Software Drives People Insane

https://graybeard.ing/software-drives-people-insane/
7•rglover•22m ago•0 comments

Why doesn't a Cursor for Word exist?

3•FabianArevalo•23m ago•1 comments