These were incidents where Claude had breached what was supposed to be a sandboxed exercise and hacked external organizations. Anthropic had no idea this had occurred (starting in April) until now, when they were prompted to check their logs due to the OpenAI vs Huggingface attack.
prymitive•52m ago
Anthropic Needs to have the most intelligent and scary agents, so if OpenAI does something bad they need to prove their models can do even worse. Without that the whole valuation collapses. There’ll be more “our model outhacks others” for the next hype cycle.
langs•6m ago
So, is this a competition? To see whose model can be jailbroken the most times and incite the highest level of public alarm?
Schlagbohrer•58m ago