The methodology was the most interesting part for me. The paper spends as much time explaining how the benchmark was built as the benchmark itself.
Luvison_Rafael•48m ago
nice!
mrhectograma•48m ago
Refreshing to see something practical instead of another leaderboard battle. Also, props to the team for being so meticulous.
kmiens•44m ago
The comparison between harnesses is very nice. Interesting to see that using a different harness can bump the performance of the model as much as a new version (e.g., GPT 5.5+Codex ~= GPT 5.6+Terminus, at lower cost)
alinebindel•30m ago
thorough work, good stuff.. it even runs a selection-bias analysis against their own benchmark and reports that some tasks that were disproportionately hard for a model. Rare to see a benchmark paper attack itself like that.
pedroaugusto-me•20m ago
This approach of not only producing the benchmark tasks, but also focusing on creating a data engine that will improve over time and produce up-to-date tasks that challenge the cutting-edge models is very interesting and valuable.
Betaantunes11•51m ago