Let's say, for the sake of argument, that the models are some multiplicative factor better on the inside.
Doesn't that mean the demos should work?
Also, yes there are still bugs in Claude Code. I experience them nearly everyday.
It is markedly better than early days, but still not the best harness.
The best software written with agents seems to come from people outside of the labs (see pi, for instance — or all of cloudflare’s recent work)
Which makes me question either the model, or the holders …
Also, not mentioned in my post:
- Mitchell Hashimoto
- Prime Intellect (and all their agent experiments)
- Geoff Huntley (see Jiti, for instance)
There's a ton of interesting software being developed with these models, but I find most of the software from these big guys to be ... bland. Buggy copies of copies.
Amid all the discussion of sigmoid curves, and where the "LLM wall" will materialise, I think few people would have predicted that the real wall in LLMs would end up being humans' capacity to verify the output.
What I fear is that people simply eschew human review altogether, considering we're talking about the industry that came up with the "move fast and break things" credo. Human review of LLM-produced code where I work is already a farce, and we're not special enough to be one of Karpathy's 5,000. I do my best to manually review anything that's my responsibility, but I'm literally one of very few people left working on my team, so in practice what happens is I submit PRs that are at best glossed over by completely unrelated teams for security, malware/prompt injection, and other serious concerns. Quality insofar as vetting others' code has completely gone out the window and it shows in the number of bug reports that come back, often themselves written in Claudease. This is all on top of everyone cynically phoning it in in the first place, due to the omnipresent sword of Damocles that is additional AI-driven layoffs.
Worse yet all the incentives point to this being the most economically viable thing individual companies can do. I think it goes without saying some type of regulation here is urgently needed, and that an unexpected cause of an AI bubble pop may end up being that humans simply aren't able to keep up with the pace of the output - leading either to precautionary plateauing of capability, or major liability risks related to a decline in quality.
That was the dominant concern in the circles I’m in, so it’s worrisome it’s being treated as rare.
I think LLMs are valuable and spend most of my professional life working with them.
I do wonder thought whether there’s a Ponzi scheme aspect here where as long as the “frontier” can keep outrunning human review and comprehension, LLMs are always going to looked way more valuable than they are and the bubble will continue.
This started with deep learning, expectations weren’t met and people started looking for value, then GPT came out and people got wooed again and forgot, then coding, then math, cyber, etc. As long as the dust doesn’t settle we never have to reflect on all the shortcomings and can just stare mesmerized at demos.
Really? I start with a conversation for maybe 5 turns or so, where I ask it what would need to change, what the API might be, any database schema changes, URL schemes, and so on, and finally ask it to break it down into commits. Then I let it go, implementing 3-10 commits at a time via subagents. It usually gets the UI somewhat wrong, so there are followups to fix it. This is with Sol and Luna subagents.
Is that what other people see?
> Somewhere around 20M people (0.2%) see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt.
No, they can't, at least not any semblance of quality. The cases we're seeing where this does kinda work is in ports and translation where all the rules are already documented in the best specification language possible with a way for the LLM to verify itself: code. We saw this close to a year ago now with Cloudflare and NextJS
> The impact scales with ambition, problem size, and horizon. A question with a paragraph answer barely stresses the system. You need a reservoir of big, difficult problems that you really care about
These are operating on different capabilities - AI's ability to answer informational queries as a chatbot frankly sucks and can't be trusted without verifying it. I run up against this every day. A problem with a verifiable answer on the other hand it's very good at solving. He knows this (his next paragraph) but he's putting them on the same scale of "stressing the system" to attempt to add proof to his introductory claim
There’s no reason to think this.
It needs to a potential field with practical, economical, "real life" attraction well. I.e, robotics, real economic efficiency gains, manufacturing novelties.
The worry is that the 1% is attracted towards a non-practical money hole. I.e: Token burn for the lols, sophisticated software systems that dont provide actual value outside of giving NVIDIA cash.
weinzierl•39m ago
The future is already here. It's just not very evenly distributed.
lifeisloving•35m ago
There is in fact no indication of this, not evem the precious benchmaxxed benchmarks ya'll love to reference.
There is however a exponential curve of slop, and an ever increasing number of peoples who's minds are completely captured by these things.
tkz1312•5m ago
gr_norm•1m ago
atmavatar•25m ago
Caveat: that same tiny group is employed by the AI vendors, meaning it's in their financial best interest to make it sound like the curve is going vertical.
njovin•18m ago
Arkhaine_kupo•15m ago
Considering the investment in AI, the lack of moat, and the increased inability of any of the big players to come even close to profitability (with OpenAI already breaking the "ads" emergency glass option)... perhaps he meant a tiny group is already seeing the line crater
nullpoint420•3m ago
AI models can dismantle billion dollar industries. They can reverse engineer Adobe and Microsoft products that once were their moats and titans of their industry.
Why do you think they'd need to lie?
albatross79•13m ago