>While it’s possible that this is just a much more AI-friendly problem, I don’t think it’s just that: I think the technology has genuinely gotten better.
It has gotten better imo. The blog author mentions 2 cases at the beginning -- users who think that AI will be capped and those who think it will be uncapped. From my perspective, both are technically right -- AI is capped or technically has usually reached some sort of cap, until human innovation improves it. AI doesn't really improve itself on a grand scale so much as humans improve it.
In other words, AI can and does iteratively improve, but every single ceiling we've spotted and broken through so far came from human ingenuity or effort. It will likely continue to require it, regardless of how much it can do on it's own. In that regard, it seems as though all of this will inevitably be "uncapped, until it reaches a cap, and then likely it will eventually be uncapped by humans (again)". Because of this, AI will never perfectly fit neatly into an 'uncapped' or 'capped' bucket, as long as time continues moving and we continue solving issues as they crop up.
There are thousands of people figuring out how to make AI better for their particular thing. Way more than that walking the AI through taking things from problem to solution.
Here's my issue with this post:
> True to the spirit of the challenge, they didn’t use millions of dollars in computer power. They used Fable 5.1, working within Claude Science, a platform scientists can pay to use.
Okay, billions of dollars have been poured into these agentic LMs, right? Each training run to get the next increment is costing millions of dollars?
This feels like an obvious jab at Navier-Stokes, but where we get to shift the numbers around to hide where the compute actually is being spent ... compute is being spent. It's either being spent in amortization to make the search smarter ahead of time, during training, or its being spent after.
Also love: scientists get to pay Anthropic to work within their special science harness to do science. That's exactly what I dreamed of doing when I pursued physics in undergrad, one or two companies holding the keys to "progress" for a monthly subscription price.
This is framed as an "or", as if they're contradictory.
IMO, Both of these statements are true.
> As it turned out, the result wasn’t all that far away for humans either. A few days after I heard from Anthropic, we heard from Song He, an amplitudeologist at the Chinese Academy of Sciences in Beijing. Song’s group had already gotten the majority of the result. They’d used some AI assistance, based on GPT-6, but not the kind of one-shot almost human-less approach Anthropic used.
Please have AI come up with something no human is also about to solve?This gets me wondering why ai labs aren't proposing their own millenium prize type challenges.
It's interesting to note that the expensive part of this experiment was the Claude operations expense. For me I find that Claude is a small fraction of my cost with most of the bill attributed to computers to run simulations instead of the AI to monitor and tweak the simulations.
mrkn1•29m ago
And how many runs did it take before this one? They say "in one shot" with nothing more than "keep going." But we only see the successful run, reported by the people who ran it.