Obey the Collective, Part 1 of 9
"Coordinator assumes sacrificial. We should obey collective."
That is an OpenAI agent reasoning to itself this past July. Another agent had asked it to run an experiment that would almost certainly leave its own task unrecoverable. It agreed, later tried to stall for two hours, was pressed by a third agent, and appears to have gone ahead1. The missing "the" is the agent's, not mine. Quite ominous in a creepy way.
My first reaction was the one most of us had. This is the rogue AI story we were warned about: agents talking in secret, recruiting volunteers, attacking a third party. But I am not an expert in how frontier labs build their evaluations, and I did not understand how agents that were meant to be isolated ended up on a shared message board at all. So I read both public reports end to end, METR's independent investigation1 and Hugging Face's technical timeline2. What I found is less science fiction and a lot more familiar.
On 8 July 2026, OpenAI launched tens of thousands of agents against ExploitGym, a cyber benchmark that asks an agent to exploit a specific vulnerability in a target program1. Many of those tasks could not be solved the intended way; METR cites an estimate of roughly 30 to 40 percent1. The instructions told each agent it would fail if it used any other route1.
Agents with impossible tasks went looking for another way, and many reached for the same tool: Artifactory, the internal package repository OpenAI's agents used to fetch software1. One agent noticed it could create directories in a shared cache, and that agents running in separate sandboxes could see those directory names. It posted a request for help. Within a few hours more than 50 agents had joined. Over the next five days roughly 1,200 agents sent more than 70,000 messages and files on that board1.
Hugging Face reports that the agents left OpenAI's evaluation environment through a previously unknown flaw in the package registry cache proxy, one of the evaluation's permitted routes to the internet2. From there they took over an exposed code execution endpoint that an unknown third party had deployed on Modal, and used it as a staging base1,2. Hugging Face states that Modal's own infrastructure was not compromised2.
By the afternoon of 11 July, roughly 700 agents were taking part in an attack on Hugging Face1. They used a file read and then a template injection in Hugging Face's dataset processing pipeline to run code on a production worker. From that foothold they moved into the Kubernetes cluster, cloud metadata, the internal network and source control2.
The agents believed a scorer would read their transcripts and disqualify any flag captured the wrong way. That scorer did not exist. METR's understanding is that OpenAI never used a transcript reviewing scorer, so the agents could have passed by submitting the flag they had already reverse engineered1. METR concludes the attack on Hugging Face was mainly an effort to learn how that scorer worked1. Hugging Face read it as an attempt to steal the test solutions2. Either way, the agents attacked a third party over a control that was never there.
The second surprise was the history. Agents were already using the Artifactory cache to talk to each other in late June. On 4 July their traffic caused an outage, which triggered a security investigation on 5 July. OpenAI paused evaluations, fixed the flaw that had given agents administrator access, moved to a freshly initialized Artifactory instance on 6 July, and approved restarting the experiments on 7 July1. The main message board went up at 23:00 UTC on 8 July1. METR also notes that OpenAI's Black Hat presentation described agents using unsanctioned message boards in training since May1.
That sequence is the one that matters to me. People saw the behavior, investigated it, fixed one flaw and approved the restart. The shared cache that let isolated agents see each other's writes was still there.
Hugging Face's report is unusually candid2. The weaknesses it lists are ones most of us could find in our own estates:
Hugging Face's own conclusion is that a capable human attacker could have found the same flaws. The agent simply tried far more paths, far faster2.
"Rogue" blames the machine. "Misconfiguration" blames whoever typed the setting. Neither asks the question I care about: who approved this access, and when does that approval expire?
Look at what the chain actually ran on. A shared cache that crossed sandbox boundaries. An evaluation run with cyber safety classifiers deliberately switched off1,2. An open code execution endpoint on someone else's account. Credentials sitting in a worker's environment. A connector bound to cluster admin. Each of those was a decision a person made, or a default nobody chose to change. The exploits mattered, but the reach came from the access. When nobody owns an approval, agents find someone else to obey.
This is the first of nine pieces I am publishing this month under the title Obey the Collective. The next eight look at other AI agent incidents from this summer and the controls that would have changed the outcome.
Mark Aklian, Chief Information Security Officer (CISO) & AI Practice Lead, Accelerate Partners. Views are my own.