1. Just variance in pass@K. If you prompt any model multiple times you'll see a large variance. N=1, but I find chinese open source models have a higher variance than higher-RL'd models like fable/opus.
2. They legitimately shipped a new RL checkpoint over the 7 days, which I find hard to believe.
I am leaning towards 1.
while here it outperforms Fable by a significant margin:
but if the latter is true, will people still say it was "distilled" from Fable?
on toy benches it made quite a few mistakes but was able to fix all of them on its own
(meaning more tokens, more turns, more tool calls — but same outcome as gpt 5.6 sol)
garo-pro•1h ago
mohsen1•6m ago
Seems legit.
It's really hard to know how good it is. So much hype around it.