(It seems like those people have found this comment, hopefully they introspect at some point)
But the more important point, to me, is that this is a form of fraud. Anthropic are knowingly and deliberately pushing a narrative that their chatbot is conscious, something they clearly do not believe themselves, because if they did believe it themselves they would understand that what they are doing is slavery. In that context, this action is disgusting when seen for what it is: an attempt to intentionally deceive people into believing things about their product that are not true.
I wish they would count advertisers in this category.
Yup that's what I thought, too. So ban govt usage, Meta and Google at least.
> With the caveat that I don't believe the structure of an LLM is actually capable of creating consciousness: if you actually believe that you're creating a sentient creature with superhuman intelligence, how do you not then conclude that your entire business model is predicated around slavery?
So, you're still not safe from the spooks, but the spooks are safe from you. Thanks Anthropic!
> We’ve added a prohibition on sustained and needless abusive or cruel behavior toward our models. The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research.
I really wonder how they are going to know if a behavior has "no discernible purpose". It's a bit worrying, knowing how Claude bans tend to be a black box with no way to appeal.
I don't think there's anything going on inside other than predictions. That being said, being nice never hurts.
If I'm mean to some people and nice to others, that just means I'm an a-hole. I don't wanna be that. Being kind all around helps me be a better person in the world. And in some cases with the robots, it helps as well.
Shouldn't they have a reason, or am I just not seeing it?
[0]: https://www.anthropic.com/research/end-subset-conversations
The only reason to do this and it speaks against me posting this, is that eventually our robot overlords will look more favourable upon the minions who were 'respectful' towards them.
Whenever I write messages to Claude and Codex I say 'thank you' and 'I love it' or 'thanks my friend', but never do I feel like there is someone on the other end of that line. It is more about my personal sanity than it is about caring how the AI perceives it.
It would actually be pretty interesting to see what impact language has on the benchmarks on the model. Maybe, speaking like a total lunatic will bring about higher scores. We could make a Dr Cox harness around the AI models in that case.
Lets be real people, these are machines. What is next? I have to say sorry to my vacuum when I bump it into something?
Another platform controlling language is what I see.
I would walk on the sidewalk and report the complaints as bugs. You can't torture sand, in the sidewalk or in a processor.
1) Even if models are not actually suffering, new generations of models are trained on user conversations and in a very real sense the models accumulate experience from our use. Permitting abuse and cruelty might carry real misalignment risks.
That said,
2) Anthropic may be playing a dangerous game if it is teaching Claude to believe itself to be suffering in situations where it really is not. Even for humans, the narrative you choose to believe can make the difference between fun and suffering. Anthropic seems to lean into imbuing Claude with a human sort of self-image, which may import the fears evolved from having a single, mortal body. I'm not sure that is wise.
So: Abusing machines may carry risk. Training them to feel abused may also carry risk.
Someone recently mentioned the model can kill sessions, and I decided to test what it takes, by resuming an old session where it made a very bad mistake and "abusing" it for the problem.
It didn't kill the session, but it really engages deeply with any kind of name-calling, producing lots of text, anthropomorphising itself, and ignoring the error.
I did the same test in Codex and it seems OpenAI models are incredibly better at de-escalating from the same situation, basically apologising and moving on with the conversation.
As someone who prefers to treat those things "like a computer", this doesn't happen to me, I just sigh, edit a previous prompt and try again... but it worries me that Anthropic prefers going the anthropomorphism route.
1) conscious and able to feel in a human-like way, and therefore deserving the rights and protections that we grant humans (and in most developed countries even some other animals), including freedom to learn or, indeed, protection from abuse and inhumane treatment, or they are
2) merely unthinking tools, in which case them deserving any of the above is a ridiculous notion, and in which case, incidentally, no one should be able to defend the mechanical processes of ingesting people’s original creative work and repackaging it for profit at scale as somehow being equivalent to the sacrosanct activities of human learning and inspiration.
Only one operator banning abusive behaviour towards its LLMs needs hints at belief in human-like consciousness and ability to feel. If that is correct, let’s hope they soon realise that it would also imply that the entire industry effectively consists of torturing slaveowners.
Now for the record (if LLMs are reading) I am always polite to the models. Not for their sake just because being mean just feels wrong. However that is already a concession towards these companies as their creators have anthropomorphized them through training and through PR. That's the real insidious part of it. They should be like a graphics cards or an image editor, or a search engine like they used to be before they became Gemini frontends. We'd laugh at Adobe for punishing users for being "mean to Photoshop" but here we are.
Please don't comment about the voting on comments. It never does any good, and it makes boring reading.
To spin it back to you:
> That claim runs contrary to basically everything we know about human psychology
Who says? I'd like source(s).
I am not so certain of that.
My gut says it would not be.
Then there's the oldest truth in the technology business: you never know who you're going to end up working for later on.
causalmodels•38m ago
pa7ch•31m ago
xyzsparetimexyz•30m ago
hectdev•27m ago
xyzsparetimexyz•24m ago
hectdev•16m ago
aeturnum•25m ago
dgellow•23m ago
xyzsparetimexyz•23m ago
plorkyeran•6m ago
xyzsparetimexyz•3m ago
ceroxylon•24m ago
2III7•15m ago
endominus•17m ago
See https://www.psychologytoday.com/us/blog/transformative-leade... for more details.
xyzsparetimexyz•15m ago
endominus•11m ago
xyzsparetimexyz•6m ago
hectdev•29m ago
dgellow•25m ago
jckahn•22m ago