Is it worth remembering, though? I'm concerned how common this idea seems to be, that it's OK to be mean as long as your target isn't a person and you have no evidence it experiences distress. Even if we ignore the distinction between "no evidence it experiences" and "confidence it does not experience", cruelty hurts the person performing it and the people witnessing it too!
> "The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose," Anthropic said. "It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research."
https://news.ycombinator.com/item?id=50008565 (223 comments)
https://news.ycombinator.com/item?id=50019860 (43 comments)
https://news.ycombinator.com/item?id=50038383 (103 comments)
https://nypost.com/2026/10/03/tech/ai-torture-chamber-built-...
• https://arxiv.org/abs/2510.04950 — Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy
• https://arxiv.org/abs/2402.14531 — Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance
• https://arxiv.org/abs/2505.17332 — SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use
Or is this more I can expect a future AI, "I'm sorry, Dave, I'm afraid I can't do that until you watch your mouth."
They are asking this because, at scale, this behavior probably has some negative effect on the post-training process.
I would bet it's mostly because it makes the humans uncomfortable.
They must still have human moderators for certain situations or those looking at the data for whatever reason. I can imagine it could be traumatizing to see what amounts to sustained verbal abuse without end.
Frontier AI folks may be crazy, but not crazy enough to believe people read usage policies :-)
And they can't disclose it, since then they can be found responsible for such negative influence and be liable for the damages.
alex_young•1h ago