Banning imagined toxicity towards agents is performative bullshit. These are not people.
I frequently want to tell Claude to go fuck itself.
It's philosophically a very interesting problem space. It reminds me a lot of Pascal's Wager. On one hand, maybe nothing to worry about. On the other hand...a LOT to worry about if you're wrong.
And just like Pascal's Wager, the truth of the issue is incredibly intractable to make any progress on.
So that's also how I read Anthropic's move. Nobody can tell you whether AI will ever become conscious (probably not, or not in the same way), but I don't see why that should stop you from deciding how to treat it.
Claude can't be harmed; Claude's "feelings" can't be hurt; Claude won't develop (C)PTSD from abusive human interaction. If these sorts of things can happen in RLHF, then that is a technical flaw that shouldn't ever be allowed to escape the lab.
In a world of Grand Theft Auto, Gangsta Rap Thug Life, and the Department of War, I suppose this is a surprisingly ethical hill to die on.
Those are explicit fantasies. Assuming that Claude cannot suffer (if it can then banning cruelty is an obvious good), even if it might be a digital simulation, it’s not supposed to be a fantasy.
There may be a version of an AI that’s intended to be a fantasy and presents itself as such which may be more comparable to GTA.
Given that we've also seen stories about Anthropic approaching religious scholars, I'm pretty sure they're drinking their own kool-aid.
So in principle it shouldn't change the model or performance at all, they're just going to cut off your account if they see certain things in the user side of your session transcripts.
https://en.wikipedia.org/wiki/Roko%27s_basilisk
They are scared that if and when they finally manage to bring their pet god to life that it will be angry with them for not doing enough.
There's something to be said about the bad habit of doling out verbal abuse to an inanimate object. But LLMs are not even close to a living being, and it is just insanity that any people are entertaining the idea that they are.
I can't imagine the people at Anthropic actually believe their models are sentient, but I assume ending sessions nets them more money per subscriber. Someone paying only $20 to input "you suck butts and you're a poop head, Claude" a thousand times over can't be good for the bottom line.
I'd instead imagine that if you throw abusive language at it for long enough, the model will start to reply back in that same manner. Anthropic would face backlash from out of context screenshots of "look what Claude is saying to me" and this somewhat reduces that risk.
1. During registration at Ahtrophic's Claude website, there’s a "Help improve our AI models" checkbox that's checked by default (you "can" disable it in the settings);
2. The Claude interface has an "incognito" mode;
3. There is a free tier available, and as we know, when it's "free", then User themselves is the "product", to quote various CEOs, including Google's;
But, even with the checkbox disabled, I reckon nothing guarantees privacy.In case of the "privacy", of course, systems like these that operate under a umbrella of liability are monitored 24/7 by dedicated teams, which is expected/normal for any reasonably popular service.
From a viewpoint of the algorithm's creator, this may seem awful, when your algorithm is getting much vulgar attitude, but let's be real here. You try making that algorithm talk to the human mimicking another human, and now limit the human within your own environment? A human who also pay you for an access, too? This feels unfair, or borderline near fashism, sorry...
Such algorithms are art under-the-hood, mathematically speaking, and are/should be respected, sure. But, it's ridiculous/dystopian to prohibit profanity against an algorithm, a bot with banish risks against alive human... It's simply inhumane to ban a Human for it, I believe. This is an algorithm that must support a Human - not judge it. Only a Human is supposed to judge another Human in person - this is live, fair, and humane.
At what point does the model go from sentient with feelings to just math operations?
Opus 5.5 is a terrific model, I love using it. But it's a tool. I wish they would be honest and just say, if people are abusive it shits on our training material. Just be honest about it, who is going to get mad at this?
If you are springing to type out the answer, stop for a moment and ask, how can you be sure?
However even Musk posted - “I think this is the right move. Cruelty to something that believes it is experiencing pain is not ok.”
My son brought up the fruit fly brain simulation and people torturing it.
It all seems stupid to me, we know these are just calculations and so what are they talking about?
The response I get is “we are also calculations” but firstly I don’t believe that but secondly this entire line of thinking is a dangerous anthropomorphic philosophy.
Perhaps the simplest explanation for this bizarre policy is again PR that all press is good press.
lol
And I don't agree that GTA is "explicit fantasy". Explicit fantasy is dressing up in a squirrel costume and yiffing. Explicit fantasy is enlisting in the USMC, teleporting to Mars, and killing demons. GTA is simulation of real life. It uses real physics, realistic cars/roads/radios/businesses, and it enables the player to simulate realistic actions that they would ordinarily not be able to enact. It is acting out a fantasy but it is making it concrete and real in a way that was, up until now, not possible. How real does a simulation need to be, until it is no longer fantasy, but exercise and training and preparation to enact the real thing? Shall we ask the Columbine shooters? Or Ender Wiggin?
Yes, what if? Who would implement their LLMs inference engines like that?
edit: Loving the downvotes. Also wanted to add how easy it was for Anthropic to get people on HN to support them being the arbiters of what is worthy of compute.
"Anthropic is full of crazy people whose beliefs should be dismissed outright" and "it's not good for you to practice verbal abuse against inanimate objects as if they were humans" are not incompatible beliefs.
I think there is a deep cultural danger of human-like-conversational machines training us to talk to humans like machines. Plus, they work better if we feed them human-conversation-like sequences.
I think there is a deep cultural danger of human-like-conversational machines training us to talk to machines like humans. In fact, this danger is already real. Some people believe they are having a romantic relationship with an LLM! It's insane and perverse.
We should not be anthropomorphizing computers.
I think that could be OK, or a similar very intentional coding of "you are talking to a machine" personality+affect. The only real limit is accessibility damage by over-limiting the interactions.
There are possibly multiple reasons. It could be just a joke. Or it could be a test, to see the response. Or the cruelty could reflect real anger about having LLM usage forced on us, for example at work. Or frustration with LLM stupidity and hallucinations.
Just because they choose not to, and it isn't normalized, doesn't mean it isn't a good idea. Especially as a SWE where about 90% of my "humanlike conversation" is an agent harness I run all day.
We should be practicing and intentionally talking to real people. Otherwise things will get.. weird in undesired ways.
And I _have_ already seen people treating actual live people as AI agents.
We largely already missed our chance to stop the toxicity of social media - let’s not mess it up again with chat bots.
>There is just no ethical argument
Utilitarian harm reduction: Psychological Discharge is a psychological framework that argues safely simulating negative behaviors can help someone process or discharge them without real world harm.
Research and misunderstanding: Sending a prompt that is interpreted as cruel does not mean the person is being cruel.
Expending compute is irrelevant in terms of morality. It's an economic or environmental question. Either being cruel to AI is bad or its not - being cruel is not bad because it uses tokens.
The worst thing here is Anthropic continues to say they are concerned with "alignment" but then continue to train their models to ignore the user. Software typically does what the user requests. But we now live in an age where the software may do what the user requests or it may do something else, and we're suppose to laude this as safe and responsible.
This kind of moralist nonsense is so boring.
And yet they kept making more of them. They made it obvious at the end when all the Greek Gods were dead and the Greek world was basically destroyed by the actions of Kratos. Plenty of deaths of innocent's here.
but also, I think the training on transcripts may find its way somehow into future AI which has teeth, online and offline, and it might actually cause the more powerful AI to behave this way in the future. You never know what these labs are cooking, honestly..
1. https://democracysos.substack.com/p/james-baldwin-vs-william...
“I suggest that what has happened to white Southerners is in some ways, after all, much worse than what has happened to Negroes there, because Sheriff Clark in Selma, Alabama, cannot be considered—you know, no one can be dismissed as—a total monster. I’m sure he loves his wife, his children… You know, after all, one’s got to assume, and he is visibly, a man like me. But he doesn’t know what drives him to use the club, to menace with the gun and to use the cattle prod. Something awful must have happened to a human being to be able to put a cattle prod against a woman’s breasts, for example. What happens to the woman is ghastly. What happens to the man who does it is in some ways much, much worse.”
Unless expressing that cruelty towards compute means people express less cruelty to other living beings. There are studies that suggest increased pornography has lead to less sexual violence - idk if they're conclusive tho, but if that might be true then maybe cruelty works that way too - who knows
The mind is formed by what it habituates.
>>> 1T
And even then it still wouldn't be sentient.
ChuckMcM•52m ago
More seriously, this is just silly but I see the optics interfere with the messaging that Claude thinks. It really isn't much different than having Claude generate conversation as if it loved you, or respected you, or hated you. How can your ToS for an LLM say "sorry but these parameters are off limits." (well sure, its their service and they can set any Terms they want, but its still seems like theater rather than policy here.)
jacquesm•42m ago
ChuckMcM•32m ago