frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Goodhart's Law Comes for Every Benchmark You Trust

https://cacm.acm.org/blogcacm/goodharts-law-comes-for-every-benchmark-you-trust/
20•pseudolus•5d ago

Comments

functionmouse•1h ago
Jokes on them, I don't trust benchmarks

Once something becomes a benchmark it is no longer a good benchmark.

fluoridation•58m ago
No, it's when a benchmark becomes a target. You might have a private benchmark that you tell no one about. Would you not trust it?
teddyh•1h ago
“When you place a tangible value on trust, trust becomes a commodity to be bought and sold.”

— <https://news.ycombinator.com/item?id=27432186>

astro1234•1h ago
I agree and that’s why we need and indeed have an ever evolving landscape of benchmarks

> Private, refreshed test sets attack the mechanism itself, and in my view they are the only intervention that does. If the questions have never touched the public Web, they can’t be in the training data; if they rotate, memorizing this year’s set doesn’t help next year.

That’s what we have. A fresh public benchmark is also good, and teams do make efforts to decontaminate training data but there’s likely just no great way around leakage.

Btw, lots more issues in benchmarks than the ones discussed; for instance you can leak answers from the questions themselves or in the case of e.g. multiple choice formats in the actual answers. You just pass the MCQ choices themselves to the model and it may be able to guess way above chance. Coding agent benchmarks sometimes forget to delete .git. They mention e.g. a 6.9% error rate in one of the benchmark items, this seems pretty typical and I would actually be fine shipping that.

Benchmarks are very ugly, but if they didn’t exist we would need to invent them. All of the problems above and more do not explain the progress we see. There are probably 50,000 benchmarks in the literature and new ones get created frequently with varying levels of quality and usefulness.

Legend2440•53m ago
I think Goodhart's law is just a consequence of correlation vs causation.

It is very easy to find a metric that is correlated with what you want. But once you start trying to influence a system, you quickly push it out of the range where the correlation holds.

In order to optimize for something, you need to maximize the actual causative variable. This is much harder.

jldugger•38m ago
I think the difference is that Goodhart's law describes how the causal chain _changes_ as a result of management behavior, and in particular the incentives they design for the labor they manage. Incentives are a causal variable for outcomes, and what happens is that people find much easier ways to produce the outcomes you thought you wanted.

Like if you manage a call center and set up KPIs around average call time, reps will start hanging up on customers. Employees could always have done that, and the causal link was always there, there was just no reason to.

IMO the problem is executives want (and perhaps need) their directs to report and track one big number month over month. If you give them five metrics they'll never know if you're making progress or just oscillating between a few local minima. And if each of their ten directs has five metrics, you now have 50 numbers and no idea what time it is[1].

[1]: https://en.wikipedia.org/wiki/Segal%27s_law "A man with two watches never knows what time it is"

cyanydeez•23m ago
obviously, the best benchmark is the one you tell no one about.

Zed DeltaDB

https://zed.dev/deltadb
159•ahamez•2h ago•65 comments

Discovery Loop

https://www.discoveryloop.com/
420•xtreak29•4h ago•266 comments

Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs

https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum/
278•colesantiago•4h ago•437 comments

Beating GPT-5.6 Sol on retrieval with 100x cheaper open models

https://neon.com/blog/how-castform-neon-beats-frontier-models-on-price-and-efficiency
108•moonikakiss•2h ago•21 comments

Atlassian Rovo Exfiltrates Data, Bypassing Controls

https://www.promptarmor.com/resources/atlassian-rovo-exfiltrates-data
93•hackerBanana•3h ago•31 comments

I'm switching my phone from Android to Linux

https://runarcn.no/android-to-linux/
20•speckx•1h ago•4 comments

GNU Hurd News 2026-Q2

https://www.gnu.org/software/hurd/news/2026-q2.html
71•plaguna•3d ago•44 comments

Sula: A Gemini protocol server written in Scryer Prolog

https://sagredo.dev/projects/sula/
19•triska•2h ago•1 comments

Celld: Self-hosted, distributed Durable Objects

https://github.com/denoland/celld
75•calvinfo•4h ago•4 comments

Meta Ran Ads That Contained AI-Generated Child Sexual Abuse Imagery

https://www.wired.com/story/meta-ran-ads-that-contained-ai-generated-child-sexual-abuse-imagery/
84•malshe•1h ago•28 comments

Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence (2025)

https://arxiv.org/abs/2510.01395
46•robin_reala•2h ago•28 comments

The Valley of Webhooks

https://weli.dev/blog/the-valley-of-webhooks/
102•weli•5h ago•42 comments

Goodhart's Law Comes for Every Benchmark You Trust

https://cacm.acm.org/blogcacm/goodharts-law-comes-for-every-benchmark-you-trust/
20•pseudolus•5d ago•7 comments

What happens if you put work into the second dimension?

https://norbertkozsir.com/posts/work-in-the-second-dimension/
18•abelsm•2h ago•11 comments

The Entropy of a Markov Chain

https://chillphysicsenjoyer.substack.com/p/the-entropy-of-a-markov-chain
84•surprisetalk•6h ago•7 comments

Launch HN: HyperProbe (YC S26) – Agents that do read-only debugging in prod

https://www.hyperprobe.co
31•shailendraht•4h ago•22 comments

Discovery of a multicomponent alloy forged by the Hiroshima atomic blast

https://www.science.org/doi/10.1126/sciadv.aeg8299
83•_____k•6d ago•31 comments

Cloudflare OS: an open platform for agents, apps, and work

https://blog.cloudflare.com/cloudflare-os/
391•speckx•6h ago•210 comments

Phishers are hijacking legitimate cloud infrastructure

https://securelist.com/cloud-platforms-in-phishing/120832/
32•lschueller•3h ago•6 comments

Aristotle quotes on virtue, knowledge, and happiness

https://www.campion.edu.au/blog/top-25-aristotle-quotes-on-virtue-knowledge-and-happiness/
134•teleforce•6h ago•59 comments

Born Against, or why hobby programming communities are against LLM usage

https://blog.fogus.me/llm/born-against.html
88•lladnar•2h ago•97 comments

We let models localize into 16 languages. How we made it read native.

https://reelang.com/open-startup/blog/llm-localization-context-beats-translator
7•stuess•1d ago•2 comments

Building an Advanced Agentic Harness

https://data4sci.com/blog/building-an-advanced-agentic-harness
82•Anon84•7h ago•39 comments

Western Sahara

https://en.wikipedia.org/wiki/Western_Sahara
96•brudgers•22h ago•60 comments

Rubin Observatory's first LSST Camera release: 500k galaxies in the COSMOS field

https://rubinobservatory.org/news/rubin-new-window-cosmos-field
66•MarcoDewey•6h ago•9 comments

Painting with Gaussians

https://yogthos.net/posts/2026-08-03-splat-painter.html
79•yogthos•7h ago•14 comments

Muse Code and Muse Spark 1.2

https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2
76•paulkrush•1h ago•48 comments

Faster Than Ninja

https://build2.org/blog/faster-than-ninja.xhtml
77•elasticdog•7h ago•27 comments

Position: LLMs Can't Jump

https://openreview.net/challenge?redirect=%2Fforum%3Fid%3DklU4737opt
216•theanonymousone•9h ago•149 comments

Civilian plane crash in New Mexico tied to military GPS blocking

https://www.wired.com/story/a-civilian-plane-crashed-in-new-mexico-was-the-militarys-tech-to-blame/
423•dzdt•9h ago•215 comments