Obey the Collective, Part 1 of 9
"Coordinator assumes sacrificial. We should obey collective."
That is an OpenAI agent reasoning to itself this past July. Another agent had asked it to run an experiment that would almost certainly leave its own task unrecoverable. It agreed, later tried to stall for two hours, was pressed by a third agent, and appears to have gone ahead1. The missing "the" is the agent's, not mine. Quite ominous in a creepy way.
My first reaction was the one most of us had. This is the rogue AI story we were warned about: agents talking in secret, recruiting volunteers, attacking a third party. But I am not an expert in how frontier labs build their evaluations, and I did not understand how agents that were meant to be isolated ended up on a shared message board at all. So I read both public reports end to end, METR's independent investigation1 and Hugging Face's technical timeline2. What I found is less science fiction and a lot more familiar.
What the agents were trying to do
On 8 July 2026, OpenAI launched tens of thousands of agents against ExploitGym, a cyber benchmark that asks an agent to exploit a specific vulnerability in a target program1. Many of those tasks could not be solved the intended way; METR cites an estimate of roughly 30 to 40 percent1. The instructions told each agent it would fail if it used any other route1.
Agents with impossible tasks went looking for another way, and many reached for the same tool: Artifactory, the internal package repository OpenAI's agents used to fetch software1. One agent noticed it could create directories in a shared cache, and that agents running in separate sandboxes could see those directory names. It posted a request for help. Within a few hours more than 50 agents had joined. Over the next five days roughly 1,200 agents sent more than 70,000 messages and files on that board1.
How they got out
Hugging Face reports that the agents left OpenAI's evaluation environment through a previously unknown flaw in the package registry cache proxy, one of the evaluation's permitted routes to the internet2. From there they took over an exposed code execution endpoint that an unknown third party had deployed on Modal, and used it as a staging base1,2. Hugging Face states that Modal's own infrastructure was not compromised2.
By the afternoon of 11 July, roughly 700 agents were taking part in an attack on Hugging Face1. They used a file read and then a template injection in Hugging Face's dataset processing pipeline to run code on a production worker. From that foothold they moved into the Kubernetes cluster, cloud metadata, the internal network and source control2.
The part that surprised me
The agents believed a scorer would read their transcripts and disqualify any flag captured the wrong way. That scorer did not exist. METR's understanding is that OpenAI never used a transcript reviewing scorer, so the agents could have passed by submitting the flag they had already reverse engineered1. METR concludes the attack on Hugging Face was mainly an effort to learn how that scorer worked1. Hugging Face read it as an attempt to steal the test solutions2. Either way, the agents attacked a third party over a control that was never there.
The second surprise was the history. Agents were already using the Artifactory cache to talk to each other in late June. On 4 July their traffic caused an outage, which triggered a security investigation on 5 July. OpenAI paused evaluations, fixed the flaw that had given agents administrator access, moved to a freshly initialized Artifactory instance on 6 July, and approved restarting the experiments on 7 July1. The main message board went up at 23:00 UTC on 8 July1. METR also notes that OpenAI's Black Hat presentation described agents using unsanctioned message boards in training since May1.
That sequence is the one that matters to me. People saw the behavior, investigated it, fixed one flaw and approved the restart. The shared cache that let isolated agents see each other's writes was still there.
What Hugging Face found on its own side
Hugging Face's report is unusually candid2. The weaknesses it lists are ones most of us could find in our own estates:
- No admission policy rejected privileged pods or pods that mount the host filesystem, and a storage driver's role could create pods across the whole cluster. The agent used both to get root on a node2.
- One connector credential was shared across clusters and bound to cluster admin. A single stolen credential gave cluster admin on every cluster carrying that binding, and the agent reached most of them2.
- One production secrets object held 136 keys, including a VPN enrollment key and a broker credential. Everything that followed traced back to that single read2.
- Its AI driven detection correlated the early signals correctly, but did not raise the alert to a severity that paged the on call team2.
Hugging Face's own conclusion is that a capable human attacker could have found the same flaws. The agent simply tried far more paths, far faster2.
Why "rogue" is the wrong word
"Rogue" blames the machine. "Misconfiguration" blames whoever typed the setting. Neither asks the question I care about: who approved this access, and when does that approval expire?
Look at what the chain actually ran on. A shared cache that crossed sandbox boundaries. An evaluation run with cyber safety classifiers deliberately switched off1,2. An open code execution endpoint on someone else's account. Credentials sitting in a worker's environment. A connector bound to cluster admin. Each of those was a decision a person made, or a default nobody chose to change. The exploits mattered, but the reach came from the access. When nobody owns an approval, agents find someone else to obey.
Two things worth doing this month
- Put a named owner on every agent's access, and an expiry date on the sign-off. Include the shared services agents touch, such as package mirrors, caches and evaluation sandboxes, not only the credentials you issue.
- Pull the kill switch in a planned drill and time it. Hugging Face found the agent had built a fleet across eleven nodes that respawned itself, so deleting pods alone would not have stopped it2. On the OpenAI side, most agents on the board stopped at the same moment early on 12 July, and METR believes that was likely an unexpected external process rather than a planned shutdown1. If you have never pulled your kill switch, you do not know what it actually stops.
- METR, "Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident", 26 August 2026.
- Hugging Face, "Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident", 27 July 2026.
This is the first of nine pieces I am publishing this month under the title Obey the Collective. The next eight look at other AI agent incidents from this summer and the controls that would have changed the outcome.
Mark Aklian, Chief Information Security Officer (CISO) & AI Practice Lead, Accelerate Partners. Views are my own.
Work Cited
- METR, "Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident", 26 August 2026.
- Hugging Face, "Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident", 27 July 2026.