frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Ask HN: Anyone interested in building a harness-only benchmark?

2•GodelNumbering•50m ago
There are a lot of LLM benchmarks but few, if any, harness benchmarks. I am thinking this would be a really good community project to build one.

End goal: a leaderboard of harness performance (multiple axis) on a set of diverse real world tasks[1], grouped by underlying models and reasoning efforts. Anyone can contribute results.

The task criteria, measurements, underlying framework et al can be decided by a group rather than a single person.

If there is sufficient interest, I will create a discord.

Disclosure: I am the maintainer of a coding agent called Dirac (https://github.com/dirac-run/dirac) so I will not influence what the final benchmark should look like to avoid any conflict of interest. I just want to make this happen.

[1] Diverse real world tasks meaning sufficiently complex tasks that the contributors have encountered, preferably from an opensource repo.

Comments

theChris-in•48m ago
Interesting idea. If you can manage the infra, I can put together a replicable test suite.
GodelNumbering•43m ago
I can manage the infra, have a lot of experience in that area. A benchmark with problems coming from multiple sources and backgrounds would be ideal
theChris-in•27m ago
We can do a mix of general use (as in user stories) plus a few academic benchmarks.

So you have any specific ideas?

You can hmu at iam@thechris.in

GodelNumbering•19m ago
Thanks, I will reach out. I have also posted for contributors on localllama https://www.reddit.com/r/LocalLLaMA/comments/1vg40w8/anyone_...

> We can do a mix of general use (as in user stories) plus a few academic benchmarks.

Yup sounds about right. Generally speaking, higher the distinct contributors, more likely it is to capture the distribution of real-life usefulness.

> So you have any specific ideas?

Only that the problems that get picked should be easy to evaluate in isolation and should test the harness capability rather than model's knowledge/capability.

theChris-in•14m ago
> should test the harness capability rather than model's knowledge/capability.

Then we need a provenance for model inference, generalized. This should be interesting. We would be trying to deterministically generalize a baseline "can do this" for models... Maybe categorize by parameter class.

How Internet Fighting Works

https://www.smbc-comics.com/comic/2013-04-07
1•fanf2•1m ago•0 comments

Kimi K3 2.78T on One CPU with 8GB RAM

https://github.com/FareedKhan-dev/kimi-k3-in-c
1•OsamaJaber•1m ago•0 comments

Show HN: Hardware requirement calculator for local LLMs

https://soverignplan.app/
1•gkrishna•2m ago•0 comments

Could psilocybin be the key to treating anorexia?

https://www.scientificamerican.com/article/psilocybin-could-kick-start-anorexia-recovery-early-re...
1•tzury•4m ago•0 comments

Show HN: An app to dump notes during the day and plan todos without distraction

https://today.spqrk.net/
1•SPQRK•5m ago•0 comments

Could a change in diet improve mental health? Metabolic psychiatry explained.

https://www.nationalgeographic.com/health/article/metabolic-psychiatry-keto-diet-brain-science
1•bookofjoe•11m ago•1 comments

Yeti Devworks

https://yetidevworks.com/
1•joshka•12m ago•0 comments

Ask HN: I built bribes.fyi, now I am stuck what to with it

1•neverenderr•12m ago•0 comments

Study: People prefer stories written by AI

https://www.cambridge.org/core/journals/judgment-and-decision-making/article/bot-or-not-can-peopl...
2•senordevnyc•15m ago•1 comments

Night-Tower: Over-Engineered Homelab Powered by Fighter Jets and AI Chaos

https://shadowfish.night-tower.net/post/night-tower-homelab-architecture/
1•taubek•15m ago•0 comments

Trump's AI protectionism has come for robotics

https://www.technologyreview.com/2026/08/03/1141056/trumps-ai-protectionism-has-come-for-robotics/
1•rbanffy•17m ago•0 comments

Kioxia's nearly faster than Optane SSD

https://www.blocksandfiles.com/flash/2026/08/03/kioxias-nearly-faster-than-optane-ssd/5282259
2•rbanffy•17m ago•0 comments

SQS consumer can hang forever by default

https://encore.dev/blog/message-queue-hangs
2•eandre•17m ago•0 comments

Drug Discovery Has No Magic Wands by Daphne Koller

https://www.a16z.news/p/drug-discovery-has-no-magic-wands
1•eamag•19m ago•0 comments

Sober: Local-first code reviewer for agentic PR (deterministic and model review)

https://sober-dev.app/
1•marmai•20m ago•0 comments

People are ghosting long-term partners. Some don't regret it

https://www.wired.com/story/people-are-ghosting-long-term-partners-some-dont-regret-it/
3•eternalreturn•20m ago•0 comments

Show HN: Free EN 16931 e-invoice validator that runs in the browser

https://einvoicekit.com/
1•yzaroui•20m ago•0 comments

Belgie – TypeScript Sandboxes and React MCP Apps for Python

https://mplemay.github.io/belgie/
3•mplemay•22m ago•0 comments

Jeremy (Snail)

https://en.wikipedia.org/wiki/Jeremy_(snail)
2•ostacke•22m ago•0 comments

Gentle Response to Dontasktoask

https://www.dontbeasillygoose.fyi/
1•yikestra•23m ago•0 comments

Why the Best Software Engineers Focus on System Design

https://www.youtube.com/watch?v=LeUUxLRdvho
2•Brysonbw•23m ago•0 comments

Show HN: Code Factory – Create a reviewable MVP in minutes, with receipts

https://github.com/zrk222/code-factory
1•zrk222•24m ago•0 comments

Simulation Apps Pinpoint Cause of Electronics Failures

https://spectrum.ieee.org/electronics-corrosion-multiphysics-simulation
1•rbanffy•25m ago•0 comments

Show HN: Nodes – A native macOS Markdown editor that runs its AI on device

https://nodes-web.com/
1•jacobchen7•25m ago•1 comments

Beyond Scarcity: How Abundance Shaped Economic Thought from Smith to Romer [pdf]

https://www.paecon.net/PAEReview/issue114/Beaudreau114.pdf
1•ike_usawa•27m ago•0 comments

The Disruption Society [pdf]

https://www.paecon.net/PAEReview/issue114/Davalos114.pdf
2•ike_usawa•28m ago•0 comments

Ask HN: When do you choose RAG over Fine-Tuning?

1•Harish_0089•29m ago•0 comments

Polling shows a growing political reckoning is coming for data centers

https://www.politico.com/news/2026/07/21/poll-data-centers-democrats-moratorium-01001799
3•mapping365•29m ago•1 comments

A graded registry of scientific results produced by or with AI

https://whataifound.org/
1•ygtisik•31m ago•0 comments

Paper Rankings across live ArXiv categories

https://kurate.org/
2•nnx•31m ago•0 comments