We can see this already with GPT how luna can do 90% of what sol is used for. There's also venues and probably the only reason why china still bothers 'distilling' models is accurate training data generation.
The more intelligent models get, the more people will offramp to cheaper solutions that get the job done. There's no real benefit to using a sota model when the accuracy is already 99% and I think that is the biggest danger to US labs.
JSR_FDED•11m ago
From the article it’s not clear how the scorer grades every episode - was it a frontier model that assigned the grade? How does that continue to work as the model that is being fine-tuned becomes better at the task than the frontier model?