Today, we're launching selfbench.dev, an open-source tool that lets you create and run evals automatically from your PRs.
Every benchmark with sufficient trust eventually gets benchmaxxed (Goodhart's law) - the labs are incentivized to maximize their scores on that benchmark, which isn't predictive on whether it'll actually work within your setup. This has been a time-consuming process that only the largest companies can afford to do, so we built self-bench to fix this.
Self-bench uses agents to author / review Harbor environments generated from your PRs. You then approve every eval that the agent created, and can model/harness evals concurrently on sandboxes.
To save you money, we also allow you to connect your OpenAI and Claude subscriptions so you don't have to pay raw token costs.
We ran evals on several large codebases like
- Posthog (https://selfbench.dev/PostHog/posthog)
- Next.js (https://selfbench.dev/vercel/next.js)
- Sentry (https://selfbench.dev/getsentry/sentry)
- Pi (https://selfbench.dev/earendil-works/pi)
- as well as other fantastic open-source projects (all on selfbench.dev)
From our evals:
- Kimi K3 and GLM 5.3 are almost always more expensive, yet less performant than models like GPT-6.1 Sol and Claude Opus 5.5, due to token efficiency.
- GPT-6 Luna is almost always the most effective "cost-efficient" model we've tested, not open-source models.
We'd love for you to try this on your codebase and let us know what you think!