Most benchmarks just test if AI can fix a problem you've already pointed out.
But obviously it would be much better to fix problems before you or any user runs into it. Like, isn't it crazy that we still have to wait for people to open tickets before a lot of obvious bugs get found?
We wanted to test that capability at scale. Turns out that models still are terrible at it (best setup we tested still fixed < 5% of bugs)
We have 100 repos of 22 languages and 4k bugs between. All the bugs are real-world bugs from github. We do a lot of filtering to ensure everything can be solved in this setting.
===========================================
Sol 5.6 (xhigh) 4.7% $7,230
Luna 5.6 (xhigh) 2.5% $224
Terra 5.6 (xhigh) 1.5% $357
Luna 5.6 (high) 1.4% $28
Opus 5 (xhigh) 1.3% $5,363
Kimi K3 0.6% $2,451
Luna 5.6 0.5% $4
GPT-5.4 Mini (high) 0.5% $122
GPT-5.4 Mini 0.2% $5
Gemini 3.5 Flash Lite 0.1% $6
===========================================
Also the best model is very expensive.
We have a lot more FAQ on the website https://swesweep.com/ Oh and we're all open-source (MIT license) at https://github.com/facebookresearch/swe-sweep
Curious what you all think!
ofirpress•48m ago