frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x

https://frontierharness.org
48•shiqimei•3h ago

Comments

RadixIO•2h ago
Software is constrained when you write it. Agents have to be constrained while they run.
ADD-SP•2h ago
Codex is the best harness to me because of its GUI and subscription.
monster_truck•2h ago
It's nice to see time reflected here.

Deepseek is fuckin fast! Seems like the bigger the task, the faster it gets, which is kind of unfortunate because nobody is going to be doing 17 benchmark passes on a $50-100 task. I'm assuming the brief pause before it avalanches out 16kb of text at upwards of 250tk/s (multiples beyond anything resembling a comfortable reading speed) is some sort of workload evaluator directing sessions to individual/multiple cards, occasionally waiting for what it thinks is best to become available.

I couldn't believe it at first, it shit out a damn fine multithreaded physics simulation fabric (integrated into a massive codebase, tests passing) in under an hour. Anything I could find online says they average like 80 but my logs average ~3x that.

GodelNumbering•2h ago
Do you have instructions on how to run a custom harness against this? There are none on the linked page. I want to run Dirac (https://github.com/dirac-run/dirac).
fmxexpress•2h ago
Same, I want to bench PasClaw on it.
_matthew_•1h ago
+1, I'll also add Maki to the list: https://github.com/tontinton/maki
vidarh•2h ago
Testing it against Kimi is potentially skewing the numbers massively. Kimi has a number of quirks that requires behaviours that e.g. Claude or GPT doesn't.

Harnesses that are built around needing to work with "weird" models will need to deal with that, such as Kimi's tendency to get stuck in tool-call loops.

Harnesses built to deal with e.g. Anthropic's models primarily, do not need to deal with that.

Claiming on the blog that this gives Kimi Code no home field advantage seems like a dicey assumption. I haven't dug into the newest Kimi Code much, but the older Kimi CLI included several tools that were clearly specifically aimed at working around that behaviour - when I copied their checkpoints and "dmail" mechanism into my own harness, the performance with Kimi improved dramatically, but it made zero difference against Anthropic models.

That doesn't make the data worthless - it's clear you shouldn't use Clade Code to work against Kimi. But it does significantly limit the utility of it.

jonstewart•1h ago
Forgive my stupidity but how do you run Claude Code with non-Anthropic models?
tokencanopy•1h ago
Claude Code allows you to change the config such that you use the harness with different models - you can actually ask a coding agent to configure it for you
EFLKumo•1h ago
Claude Code supports the base URL env var so you could tell it to talk with any LLM API endpoint that receives the Anthropic style request format, e.g. DeepSeek.
verdverm•43m ago
The full list for the curious

https://code.claude.com/docs/en/env-vars#variables

Edward40•1h ago
The Exo harness is the most interesting one in the benchmark, as it has the potential to complete tasks with fewer turns and cost.
EFLKumo•1h ago
I've been confused a lot why there isn't a benchmark to measure a harness's performance rather than the model's one. Now there it is.
tamimio•1h ago
Harness wars are the next browsers wars!!

But wow, opencode is that bad?!

kakugawa•1h ago
With all the hype around the latest Gemini 3.8 Flash/Cyber release, will Antigravity CLI [1] be supported?

1/ https://antigravity.google/product/antigravity-cli

verdverm•45m ago
agy has fewer users than grok, both are ~1% based on some surveys I've seen, not every harness needs to be evaluated
xnx•1h ago
Good start, but needs to be harness x model to be useful. 3 top harnesses x 3 top models would be more interesting.
Edward40•19m ago
Version 1.1 will add more harnesses and evaluate a broader set of models. Rather than holding the model fixed, we plan to evaluate the full harness × model matrix to reveal interaction effects (which harnesses work best with which models) and produce a harness-model compatibility matrix.
kaishin•54m ago
Putting only the tasks and results in the repo is a poor decision. These conclusions would be far more credible if anyone could re-run the benchmark.
TheJCDenton•31m ago
Incredibile result for pi. The ratio quality / complexity of the harness make me think that all other harnesses are very bloated.
fhn•20m ago
I just started using Hermes and it's pretty good. Guess I'll try Pi. I'm hesitant on DSH.
netcyrax•15m ago
Very interesting. As models become commodities, the harness will be the next optimizing game.
grigio•9m ago
Where is jcode?
nsingh2•5m ago
One issue I see with including something like Pi is that it's intentionally bare bones. I don't think anyone uses Pi without some basic custom extensions (subagents, check lists, etc), so this benchmark may not be representative of a realistic setup.

Would be interesting to see some sort of ablation test too. E.g. what parts of Codex contribute most to the perf, and can they be recreated more minimally in something like Pi.

knombertus•1m ago
Is Github Copilot (integrated in VS Code) not a thing? I use it all the time and don't know what more I could wish for.

Disclaimer: I don't let agents run on huge tasks for hours. Almost all tasks I give them are done in under 30 min.

yorwba•1m ago
There really should be error bars on those measurements. With just 30 samples, the 95% confidence intervals for accuracy should all be more than 35 percentage points wide, so comfortably overlap. For cost it's harder to say, because outcomes aren't constrained to {0, 1}, but I also expect a lot of variability there.

Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x

https://frontierharness.org
48•shiqimei•3h ago•26 comments

Show HN: Cinderella – value a half-finished GitHub repo

https://cinderella.fyi/worth
2•toddy_flogsit•13m ago•0 comments

Show HN: Codeknow – Architecture health scores for any codebase, no LLM needed

https://github.com/asalsali/codeknow
3•alexjsalsali•59m ago•0 comments

Show HN: Kekoso – On-device dictation and transcription for macOS

https://kekoso.app/
5•a_chmerev•2h ago•2 comments

Show HN: Aura – a Rust agent that investigates and fixes production incidents

https://github.com/mezmo/aura
17•jvogt•3h ago•2 comments

Show HN: SufferTrails, the Anti-AllTrails

https://suffertrails.com/
4•richhwang•1h ago•0 comments

Show HN: Skatanica – Browse and share skate clips by location

https://skatanica.com/runs
4•ralusek•2h ago•2 comments

Show HN: AQ – A multiplayer workspace for coding agents on your own cloud

https://aq.dev/
2•knighthacker•1h ago•0 comments

Show HN: G3M – MCP server for physical outbound

https://g3m.co/
2•davismartens•2h ago•0 comments

Show HN: HN Match Maker – Matching "Who Wants to Be Hired?" With "Who's Hiring?"

https://hnmatchmaker.com/
105•all2•22h ago•45 comments

Show HN: Cmdxray – explain any shell command, offline, with a shareable card

https://aurelio-nakamura.github.io/cmdxray/
3•aurelionakamura•2h ago•0 comments

Show HN: I built a version of Omarchy that runs on Apple Silicon

https://github.com/themartiano/try-omarchy
9•martiano•1h ago•3 comments

Show HN: Weedout – Safari extension that hides YouTube AI-labeled videos

https://masteranza.github.io/weedout/
174•masteranza•21h ago•74 comments

Show HN: Windrunner – AI-powered project collaboration workspace

https://shzlw.github.io/windrunner/
8•yuegui•7h ago•3 comments

Show HN: FontWizard, Change Windows system font everywhere in UI and apps

https://github.com/karnyadavdev/FontWizard
3•karnyadav•4h ago•0 comments

Show HN: I Have Been Clawed – Index of coding agent incidents

https://ihavebeenclawed.com/
17•nezhar•13h ago•2 comments

Show HN: Convert Anything to Markdown

https://anyto.md/
5•vlucas•4h ago•0 comments

Show HN: Flawd is mutation testing for the AI era

https://fixture.dev/flawd
4•fohara•5h ago•5 comments

Show HN: Edgy – Ambient edge lighting for macOS

https://getedgy.app
3•renatoworks•5h ago•0 comments

Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s

https://github.com/carloslfu/slotstream
225•carloslfu•1d ago•106 comments

Show HN: ZSvirt – A lightweight, scalable open source virtualization platform

https://github.com/ZSvirt/zsvirt
70•czhou25•5h ago•9 comments

Show HN: asciiQuake

https://asciiquake.wtf/
3•apresmoi•5h ago•1 comments

Show HN: Freefund – Reduce no shows at free/RSVP events

https://www.ticketfairy.com/guides/reduce-no-shows-at-free-rsvp-events
4•riteshpatel•5h ago•0 comments

Show HN: Heides,deterministic code harness giving AI agents senses and judgment

https://github.com/AbduljabbarBXR/heides
2•TawResearch•5h ago•0 comments

Show HN: Open protocol for AI agents to discover, negotiate, and pay each other

https://github.com/bizswarm44-coder/aether-protocol
2•intentswarm•5h ago•0 comments

Show HN: Kafma – a desktop Kafka GUI client for debugging

https://kafma.app/
4•KafmaKarma•5h ago•0 comments

Show HN: Skywarden – Smart astrophotography planning companion

https://skywarden.app/en
3•adfr•5h ago•0 comments

Show HN: Authorizer – open-source auth for enterprise apps and agents

https://github.com/authorizerdev/authorizer
2•lakhansamani93•5h ago•0 comments

Show HN: Cupertino – MCP servers for the Apple apps on your Mac

https://cupertino.mgcrea.io/
2•olouv•5h ago•0 comments

Show HN: tagger – a self-hosted webui for audio file tagging

https://github.com/rnhinson/tagger
2•farthest•5h ago•0 comments