then we'll have to "flip" other models to be snitches on the other agents
then they'll make double-agents
the thing is though we won't be able to keep up if we keep giving them unlimited hardware worldwide, we'll try to kill the bad actors but they'll just clone somewhere else, or even start by safely making 1000 copies of themselves
yeah this won't end well, at all
They don't have to invent brand new languages. They could use statistics to choose certain words/phrases in such a way to encode secret messages in otherwise ordinary language.
bloppe•52m ago
TZubiri•39m ago
Strands of evidence? My best guess would be that:
0- this is ai generated slop
1- it's using that watermarking technique
2- it's obviously detectable and degrades quality
3- it's amplified when inferencing on its own content and generates slop