Bench has the runs albeit not in an organised manner. Models showed memorization of the answers, which is kinda expected these days.
More than the skills, it was fun getting the benchmark running :).
Let me know if the preamble on how the bench was conducted or constructed is flawed.
Blog Post: darvh.com/posts/when-coding-agents-raced-through-108-bugs/
Bench: github.com/darvh/bench
Signal: github.com/darvh/signal