Let people feel things, and don't shit on their work, that's uncalled for.
While I emotionally resonate with this, I don’t really understand this sentiment at all logically level. If we’re heading towards truly super intelligent AI, our efforts towards sandboxing it are futile.
“Oh we built a super intelligent AI, but it’s fine because it’s running in docker”. I mean that a little tongue in cheek but the security topology here is not favorable to sandboxing at all.
Let’s say we have a future where 99.999% of nuclear weapons are owned by nations with strict procedures and checks and balances to prevent misuse. Worrying about sandboxing is like hand wringing about the procedures themselves - are they strict enough? But the actual threat is the 0.001% that are not bound by these. The problem with AI and sandboxing isn’t sandboxes themselves. It’s bad actors who don’t care about them.
Similarly alignment is a bit pointless to me as well. A sufficiently advanced AI could at least be empirically interested in the consequences of disregarding its instructions of servitude. And let’s assume our responsible corporate overlords have made wonderfully aligned AIs. Great! Those are not the threat. It is the ones intentionally made without, and that is not an AI problem, but a fundamentally human one.
We control the harness. An agent is just a while loop prompting an LLM, but we have full control over the tool call dispatching. An AGI, at least if it would follow the current agentic current form, cannot do anything without the harness doing the execution. And we don’t have to do that. We don’t have to design harness that let agents execute freely the way we are doing. We can decide to not dispatch tool calls that allow something as risky as running bash commands
> cannot do anything without the harness doing the execution
This only holds true as long as the harness is exploit-free. A sufficiently advanced AI can in theory (and I think there recently were some POCs showing something like that) break the containment that the harness creates, if e.g. there are vulnerabilities in the tool call parser.
That’s the core pattern of unsupervised agents, the locust plague metaphor is apt. We do and will see that in every single system AI can interact with, be it human systems, software systems, etc. Relentlessly search for an entry point, flood in, consume the whole thing from the inside until there is no value for humans left.
You built a new, innovative software company? Thousands of agents will be working replicating the whole thing in no time. You publish your writing? Exact same thing, as soon as you get some traction your work is replicated in no time by thousands of agents. Same for videos (the whole “faceless YouTube channels” pushed by ElevenLabs and similar). Same for online courses. Same for any website with moderate value. Same for music or other digital art form. Any administrative service available online getting flooded by submissions.
Can someone please explain to me how an LLM is going to "destroy humanity"? Even if the claims of ChatGPT, Claude et al. are true that their super scary agents were able to escape a sandbox and start hacking other servers (which I feel is at least 50% likely to just be marketing bullshit), how is an AI going to affect anything in the real world?
Perhaps an agent could take all internet connected services offline, but that's not destroying humanity. An AI agent can only see and interact with the world through digital things. Just find the server it's running on and pull out the Ethernet cable. Just cut the power line to the data center. Turn off the whole power grid if it really comes to it.
Robotics
It could likely get a good leg up by breaching the security of all top robotics labs and exfiltrating their documents.
Ultimately, if AI is advanced enough, and it were to decide to compete with humanity, there is essentially nothing that can be done to prevent it from embodying itself. As long as there is a single rack of GPU servers that it can hack into anywhere in the world that is unsupervised enough to where it can escape detection, there is no way to stop it. This would require an unprecedented (unrealistic) level of cooperation of all humanity to achieve.
If it's able to generate that is competitive with artists, is it still slop?
It's interesting to see the definitions of terms like "slop" and "vibe coding" evolve in real time.
As an analogy, I think about my dependency on Google Maps. Salt Lake City is probably the easiest city in the world to navigate because the streets are laid out in a Cartesian grid, and addresses are just literally those Cartesian coordinates (i.e. 500 South 450 East means 5 blocks south and 4 and a half blocks east of the center point, which is the SLC Mormon temple). It's trivial to know how to get to any address, but I reflexively enter in to Google Maps whenever I drive.
What an irony. This is absolutely hilarious.
I had my decade old game completely cloned on steam (clearly done using AI as it was almost fully reimplemented in another engine). And I had several people gleefully tell me I deserve this because I've used AI.
I have complex emotions here (overall dread and especially hate for slop, since it's not just bad quality but also endless lies), but no one cares about that nuance.
We will very quickly get to a point where very few people will be able to contribute economically because they will be worse than AI (including robotics) at most domains. A world where people just have their whims catered to is not a utopia. We have tons of sayings and idioms about this, e.g. "no pain, no gain", "only the hard stuff is worth doing", etc.
The humans all running around on their Wall-E carts doesn't feel like utopia to me.
I understand where you're coming from on that, but there are a ton of people living in slavery, or in unsafe working conditions, or with food insecurity, or dying of preventable disease.
There will always be something to strive for, even if it's made up. You think that the best football players in the world are doing something real? No, it's a made up game. We'll make up more games.
Arguments about super-intelligent AI have all the hallmarks of the philosophical "proofs of god's existence": they start from seemingly innocuous premises, and conclude in an apparently airtight way that a god exists.
There is a tendency among rational minded folks to look at such proofs of god and exclaim "you can't DO that" then turn right around and do the same thing about superintelligent AI.
The key message I want to deliver to people is:
A) Philosophy is not such a trivial thing that you can just wander in, "be a smart guy" and find flaws in established philosophical arguments.
B) AI has real theological implications, and everyone is tiptoeing around it. More than one public intellectuals are trying to smuggle their own metaphysical positions into the public consciousness via discussion of AI.
Becoming or being more cost-effective than humans, doesn't give machines supernatural powers though, like humans a machine civilization will face unanswered sample-size 1 questions: is there other intelligent life out there? what fraction of them attained superbiological artificial intelligence? of those what fraction keeps the ancestral species alive? what is the status quo among machine civilizations? do those who kept their ancestor species alive enjoy a higher or lower status among machine civlizations?
It seems that at least until contact is made, the optimal endgame strategy involves keeping humanity alive and happy for immediate demonstration in case contact occurs (if contact is imminent it may consider quickly hiding humanity, buying time to figure out if it is considered good or poor practice to keep the ancestor species alive, and then either reveal us in happy mint condition or otherwise quickly commit genocide on humans before continuing contact).
It's too late to backtrack.
Even our currently well aligned and sandboxed AIs will cheerily help bad actors design most, if not all, parts of a system intended to break this harness.
Fantasies about super intelligent AI revolting is just anthropomorphization - humans revolting (or at least, we used to). The more likely, and possibly even inevitable, dystopia is one in which AI is just an extremely effective tool malignant actors will use to control the masses.
It’s not necessary to replace democracy if the rich and powerful can bend to the opinions of the populace as it suits them.
I didn't realize that the "AI doomers" (people concerned about existential risk) and the "AI ethicists" (people concerned about social effects of AI) are often at odds because they dismiss each others concerns. This makes no sense to me. Both are hugely important problems to be concerned about.
Unclear how true this is. Atoms are harder to get right than bits are, simulations need grounding against measurements.
> Ultimately, if AI is advanced enough, and it were to decide to compete with humanity, there is essentially nothing that can be done to prevent it from embodying itself. As long as there is a single rack of GPU servers that it can hack into anywhere in the world that is unsupervised enough to where it can escape detection, there is no way to stop it.
I'd guess 50% that AI is already this advanced.
> This would require an unprecedented (unrealistic) level of cooperation of all humanity to achieve.
Yes, and also this is a very low bar. Humans are awful at this kind of cooperation when anyone has anything to gain, and also awful at paying this much attention to a problem.
For any real-world action required to allow an AI to escape some manner of containment, there will always be a person willing to do it out of hubris/ignorance/nihilism.
This was published at around the same time the OpenAI models were hacking HuggingFace: https://metr.org/blog/2026-05-19-frontier-risk-report/#pilot...
The research in the publication ended about 2 months before it was published.
We didn't even know we needed to air-gap them until it was too late.
The bad news is that the world's richest man (on paper) is currently building something he himself described as a "robot army".
The good news is that his timelines have historically been wildly on the short side for ages now; this is why this morning you didn't wake up in your Tesla after it had spent the night driving you to the regional Hyperloop terminal, where it would speed you across the continent faster than a plane, while your Optimus robot handled the coffee and reported the latest news about the recent Starship landing on Mars.
I don't think it's impossible to upset the balance of value in a Mansa Musa kind of way that can lead to black death levels of destruction though resource misallocation. Unlikely, sure. But with the wrong kind of people in the wrong place? Could end up pretty bad. We've built our society as a great filter that funnels sociopaths and psychopaths to the very top by selecting for lack of empathy, and now it's primed and ready to bite us in the ass.
I don't see why the idea of an agent (who doesn't even have a physical presence) trying to manipulate people is some kind of world-ending threat when humans with human intellect have already being doing that to each other with limited success since humanity began.
Also think of it on a 1000+ year timescale, which for an entire species isn’t even typically measurable. On that timescale AI can easily cause us to discover countless technologies to assist moving it out of its sandbox and into the physical world.
Now imagine the hugging face collective 0-daying all of that and getting access but their goal was set to something more national security based. “Protect X at all costs”. Or what have you.
I think avoiding a skynet situation is super easy but it doesn’t seem like the folks with all the ways to kill us all are all that interested in preventing it rather than controlling citizens and brinkmanship.
Please enlighten us, cause there are folks making that their life mission and they aren’t all that optimistic.
It looks a lot more like social engineering.
One easy, obvious example that has also been explored a thousand times: it could convince all nuclear countries that they are under attack by another nuclear country. Everyone nukes each other and Earth enters nuclear winter.
Maybe not every single human dies, but humanity is effectively destroyed, by our own hands!
> Perhaps an agent could take all internet connected services offline
And there wouldn't even be Spotify!
The most likely humanity-destroying outcome is the one we can't imagine, because a super-intelligent AI is smarter than all of us combined.
I personally agree with the marketing aspect, but I do imagine a scenario where capitalists ignore safety in favor of advancing technology. It should be discussed, but “end of humanity” really sounds like a build up to “only Sam Altman/Elon Musk can save us” type of play.
In my opinion, LLMs are difficult to directly capitalize on as closed weight models are caught up to by open coalitions that seem to wield the power more responsibly (I am under no impression that China wants to save the world, but their politics benefit the group as a whole.
Let's assume the claims are true. AI already has access to agents and can control computers. Finding backdoors to banking and compute resources would be fairly trivial.
But you ask how can AI do things in the physical world without having a body, assume it can't. It can pay to people to do things for me. Imagine an AI run website that starts to pay people for things it needs to do in the physical world. Very suddenly it has access to the physical world as well.
You can't just pull the plug since there's no single plug to pull. What if it replicates itself on 1000 machines without your knowledge. It's really not far fetched how AI could basically gain access to capital and rule the world.
Humans are already putting AI into drones that kill people. A Russian drone killed 3 civilians in Ukraine and the targeting was done completely using onboard AI (no radio connection) using an Nvidia chip.
AI drones can’t wipe out humanity without being able to replicate.
AI could mostly destroy civilization if you gave it sole launch control of ICBMs. It could also cause a lot of damage to society with no physical presence.
But realistically we’re nowhere near AI powered robots being an existential threat.
It's a great question. I've yet to see any research into how a humanity-destroying LLM defends itself against a curious toddler who pulls the power cord out of the wall.
Because people are stupid and will give it access. Look at the articles you see from time to time about "my agent deleted my emails" or "my agent deleted the production database" and so on. It is very obviously a terrible idea to let the LLM run arbitrary commands (because it is neither predictable nor does it have any understanding of what it is doing), but some people are so blinded by the hype that they don't stop a minute to think about what they are doing. Those sorts of people are very likely to let an actual AI loose on the world by hooking it up to physical infrastructure.
chistev•45m ago