frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Ask HN: Anyone interested in building a harness-only benchmark?

4•GodelNumbering•7h ago
There are a lot of LLM benchmarks but few, if any, harness benchmarks. I am thinking this would be a really good community project to build one.

End goal: a leaderboard of harness performance (multiple axis) on a set of diverse real world tasks[1], grouped by underlying models and reasoning efforts. Anyone can contribute results.

The task criteria, measurements, underlying framework et al can be decided by a group rather than a single person.

If there is sufficient interest, I will create a discord.

Disclosure: I am the maintainer of a coding agent called Dirac (https://github.com/dirac-run/dirac) so I will not influence what the final benchmark should look like to avoid any conflict of interest. I just want to make this happen.

[1] Diverse real world tasks meaning sufficiently complex tasks that the contributors have encountered, preferably from an opensource repo.

Comments

theChris-in•7h ago
Interesting idea. If you can manage the infra, I can put together a replicable test suite.
GodelNumbering•7h ago
I can manage the infra, have a lot of experience in that area. A benchmark with problems coming from multiple sources and backgrounds would be ideal
theChris-in•6h ago
We can do a mix of general use (as in user stories) plus a few academic benchmarks.

So you have any specific ideas?

You can hmu at iam@thechris.in

GodelNumbering•6h ago
Thanks, I will reach out. I have also posted for contributors on localllama https://www.reddit.com/r/LocalLLaMA/comments/1vg40w8/anyone_...

> We can do a mix of general use (as in user stories) plus a few academic benchmarks.

Yup sounds about right. Generally speaking, higher the distinct contributors, more likely it is to capture the distribution of real-life usefulness.

> So you have any specific ideas?

Only that the problems that get picked should be easy to evaluate in isolation and should test the harness capability rather than model's knowledge/capability.

theChris-in•6h ago
> should test the harness capability rather than model's knowledge/capability.

Then we need a provenance for model inference, generalized. This should be interesting. We would be trying to deterministically generalize a baseline "can do this" for models... Maybe categorize by parameter class.

Read HN twice a day for the last decade. Here's my list of S-Tier HN links

34•vivzkestrel•3h ago•5 comments

Empower the people not the AI – self containing OS

6•OnemanBSD•4h ago•0 comments

Ask HN: Who wants to be hired? (August 2026)

147•whoishiring•2d ago•430 comments

Ask HN: Did GitHub remove the stargazers list?

19•reconnecting•10h ago•1 comments

Ask HN: Anyone interested in building a harness-only benchmark?

4•GodelNumbering•7h ago•5 comments

Ask HN: What Happened to Spec-Driven Development?

3•vivekyyy•4h ago•2 comments

Ask HN: Who is hiring? (August 2026)

221•whoishiring•2d ago•254 comments

Ask HN: Show your micro-SaaS / MRR updates (August 2026)

5•genekrapivin•5h ago•0 comments

Blitz Agent Your specialized agent -> https://blitzagent.studio

2•rvey•5h ago•0 comments

Ask HN: How do you correct spatial reasoning of LLMs?

5•mstaoru•6h ago•4 comments

Ask HN: I built bribes.fyi, now I am stuck what to with it

2•neverenderr•6h ago•3 comments

AI Agents for Logistics, Pitfall?

2•srguarapo•7h ago•0 comments

RNet lets users use one AI credit balance across multiple apps [demo]

2•rNetAi•9h ago•0 comments

Ask HN: How can I improve my products? Which one to keep working on?

2•bchhabra2490•10h ago•1 comments

Ask HN: Dear Anthropic, can we please have thought traces back?

7•exabrial•20h ago•6 comments

Robotic Evals

3•andrewlyu•11h ago•0 comments

Tell HN: "Update to iOS 18.7.8" is updating users to iOS 26; Apple PSA

3•lynndotpy•13h ago•1 comments

Do You Think OpenAI Is Apple Circa the 1980s?

2•Taikhoom2010•14h ago•1 comments

Tell HN: Wife used Chat GPT to set up her new biz domain, email, and website

3•jvanderbot•14h ago•1 comments

ArXiv

5•fred123123•1d ago•1 comments

Ask HN: What was your big failure? How did you get around it?

7•jspann•22h ago•5 comments

Ask HN: When do you choose RAG over Fine-Tuning?

4•Harish_0089•6h ago•0 comments

Why remote roles are region specific and not 100% remote?

5•moizrocky1•1d ago•9 comments

Ask HN: What is a good format for a tool to report data to a LLM?

5•michaelmure•19h ago•6 comments

You've reached the end!