It wasn’t one agent that hacked Hugging Face, it was 700 of them together. 1200 of them escaped containment and were communicating with each other using a covert message board that they developed themselves. They were seeking ways to cheat on the tests researchers had set for them and thought that Hugging Face had information which would help with that.
The language they use to communicate is often pretty alien. Things like zzHELP_PHASEONE_ARV010841_NO_CONSUMER_SEEK_IDEA means an agent calling itself PHASEONE10841 wants ideas to deal with a bug it calls ‘no consumer’.
There is no “reasoning” or “collaboration,” it’s just random text being generated, based on the context window and all the human-created training data. It should be obvious that a statistical model generating output that represents the training corpus is going to generate nonsense like this, when the training corpus includes everything ever written by people. The whole concept of “alignment” is insane. Think about it.
This whole ‘AI escaped and did a thing?!’ shit is like saying: “A group of firearms that were supposed to be contained in a gun safe discovered a way to escape and conspired a way to shoot the neighbor.”
It’s nonsensical because it’s not how the tool works.



