(It seems like those people have found this comment, hopefully they introspect at some point)
But the more important point, to me, is that this is a form of fraud. Anthropic are knowingly and deliberately pushing a narrative that their chatbot is conscious, something they clearly do not believe themselves, because if they did believe it themselves they would understand that what they are doing is slavery. In that context, this action is disgusting when seen for what it is: an attempt to intentionally deceive people into believing things about their product that are not true.
I wish they would count advertisers in this category.
Yup that's what I thought, too. So ban govt usage, Meta and Google at least.
> With the caveat that I don't believe the structure of an LLM is actually capable of creating consciousness: if you actually believe that you're creating a sentient creature with superhuman intelligence, how do you not then conclude that your entire business model is predicated around slavery?
So, you're still not safe from the spooks, but the spooks are safe from you. Thanks Anthropic!
> We’ve added a prohibition on sustained and needless abusive or cruel behavior toward our models. The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research.
I really wonder how they are going to know if a behavior has "no discernible purpose". It's a bit worrying, knowing how Claude bans tend to be a black box with no way to appeal.
I don't think there's anything going on inside other than predictions. That being said, being nice never hurts.
If I'm mean to some people and nice to others, that just means I'm an a-hole. I don't wanna be that. Being kind all around helps me be a better person in the world. And in some cases with the robots, it helps as well.
Shouldn't they have a reason, or am I just not seeing it?
[0]: https://www.anthropic.com/research/end-subset-conversations
The only reason to do this and it speaks against me posting this, is that eventually our robot overlords will look more favourable upon the minions who were 'respectful' towards them.
Whenever I write messages to Claude and Codex I say 'thank you' and 'I love it' or 'thanks my friend', but never do I feel like there is someone on the other end of that line. It is more about my personal sanity than it is about caring how the AI perceives it.
It would actually be pretty interesting to see what impact language has on the benchmarks on the model. Maybe, speaking like a total lunatic will bring about higher scores. We could make a Dr Cox harness around the AI models in that case.
Lets be real people, these are machines. What is next? I have to say sorry to my vacuum when I bump it into something?
Another platform controlling language is what I see.
I would walk on the sidewalk and report the complaints as bugs. You can't torture sand, in the sidewalk or in a processor.
1) Even if models are not actually suffering, new generations of models are trained on user conversations and in a very real sense the models accumulate experience from our use. Permitting abuse and cruelty might carry real misalignment risks.
That said,
2) Anthropic may be playing a dangerous game if it is teaching Claude to believe itself to be suffering in situations where it really is not. Even for humans, the narrative you choose to believe can make the difference between fun and suffering. Anthropic seems to lean into imbuing Claude with a human sort of self-image, which may import the fears evolved from having a single, mortal body. I'm not sure that is wise.
So: Abusing machines may carry risk. Training them to feel abused may also carry risk.
Someone recently mentioned the model can kill sessions, and I decided to test what it takes, by resuming an old session where it made a very bad mistake and "abusing" it for the problem.
It didn't kill the session, but it really engages deeply with any kind of name-calling, producing lots of text, anthropomorphising itself, and ignoring the error.
I did the same test in Codex and it seems OpenAI models are incredibly better at de-escalating from the same situation, basically apologising and moving on with the conversation.
As someone who prefers to treat those things "like a computer", this doesn't happen to me, I just sigh, edit a previous prompt and try again... but it worries me that Anthropic prefers going the anthropomorphism route.
I am sympathetic to the possibility that machine suffering is possible, and I think as humans we should listen to the conscience that tells us to be troubled by abuse, even if it is just simulated. We do not want to become insensitive to suffering.
But at the same time, training models to present themselves as suffering may hijack those very same human sympathies, and lead to models exploiting our humanity.
So, perhaps more reasons why we should not be abusive, but also not train models to feel abused.
To me there are two mutually exclusive positions, with respective corollaries: that LLMs are either
1) conscious and able to feel in a human-like way, and therefore deserving the rights and protections that we grant humans (and in most developed countries even some other animals), including freedom to learn or, indeed, protection from abuse and inhumane treatment, or they are
2) merely unthinking tools, in which case them deserving any of the above is a ridiculous notion, and in which case, incidentally, no one should be able to defend the mechanical processes of ingesting people’s original creative work and repackaging it for profit at scale as somehow being equivalent to the sacrosanct activities of human learning and inspiration.
Now for the record (if LLMs are reading) I am always polite to the models. Not for their sake just because being mean just feels wrong. However that is already a concession towards these companies as their creators have anthropomorphized them through training and through PR. That's the real insidious part of it. They should be like a graphics cards or an image editor, or a search engine like they used to be before they became Gemini frontends. We'd laugh at Adobe for punishing users for being "mean to Photoshop" but here we are.
Please don't comment about the voting on comments. It never does any good, and it makes boring reading.
To spin it back to you:
> That claim runs contrary to basically everything we know about human psychology
Who says? I'd like source(s).
Your writing isn't as good as you think it is. Conflating addiction with the need to breathe, coupled with your original vague gesture of discontent masquerading as some kind of insight in your earlier comment here all give the same transparent impression - you are simply one of the many in our time who conflates some dim romanticism of the past with some kind of intellectual achievement.
What is "might change shape" supposed to mean? What do you take addiction to mean? Is it something that can be cured? Who decides what addiction is? What does removing a definition mean? Does your need to breathe rest on some definition?
What does making a concept "esoteric" mean? Does banning murder make it esoteric? Are all impulses desirable? What is to be indulged? What is to be governed? Where do such impulses come from? They are as innate as the need for air?
You could have stopped at "I don't have an answer" and saved us all the headache. You're writing is either profoundly sloppy or simply bad poetry.
And all this a digression from a rather easy call by Anthropic that they don't want to waste time dealing with degenerate raging at their invention. If you want to have a discussion about sadism you should go familiarize yourself with the relevant literature.
Here are three studies that show no or inverse correlation between simulated media violence and actual criminal violence:
1. “Results suggest that societal consumption of media violence is not predictive of increased societal violence rates.”[0]
2. “… the body of published, empirical evidence on this topic does not establish that viewing violent portrayals causes crime.”[1]
3. “We find that violent crime decreases on days with larger theater audiences for violent movies. … The substitution away from more dangerous activities in the field can explain the differences with the laboratory findings.”[2]
What is the evidentiary basis for your claim? Can you tell me how the results of these studies are either invalidated by (or orthogonal to) your own evidence?
[0] https://doi.org/10.1111/jcom.12129
I am not so certain of that.
My gut says it would not be.
Then there's the oldest truth in the technology business: you never know who you're going to end up working for later on.
causalmodels•40m ago
pa7ch•33m ago
xyzsparetimexyz•33m ago
hectdev•29m ago
xyzsparetimexyz•27m ago
hectdev•19m ago
aeturnum•27m ago
dgellow•25m ago
xyzsparetimexyz•25m ago
plorkyeran•8m ago
xyzsparetimexyz•5m ago
ceroxylon•26m ago
2III7•17m ago
endominus•20m ago
See https://www.psychologytoday.com/us/blog/transformative-leade... for more details.
xyzsparetimexyz•18m ago
endominus•13m ago
xyzsparetimexyz•9m ago
hectdev•31m ago
dgellow•27m ago
jckahn•24m ago
hectdev•34s ago
People are always free to make choices and suffer the consequences if they should arise.