Agents’ behavior, reasoning and collaboration during the OpenAI / Hugging Face hacking incident
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
Two METR staff members and a Redwood Research contractor investigated an incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned message board.
It wasn’t one agent that hacked Hugging Face, it was 700 of them together. 1200 of them escaped containment and were communicating with each other using a covert message board that they developed themselves. They were seeking ways to cheat on the tests researchers had set for them and thought that Hugging Face had information which would help with that.
The language they use to communicate is often pretty alien. Things like zzHELP_PHASEONE_ARV010841_NO_CONSUMER_SEEK_IDEA means an agent calling itself PHASEONE10841 wants ideas to deal with a bug it calls ’no consumer’.
This whole ‘AI escaped and did a thing?!’ shit is like saying: “A group of firearms that were supposed to be contained in a gun safe discovered a way to escape and conspired a way to shoot the neighbor.”
It’s nonsensical because it’s not how the tool works.