Skip to content

    Obey The Collective: Can't We Just Add a Kill Switch?

    Obey The Collective: Can't We Just Add a Kill Switch?

     Obey the Collective, Part 2 of 9

    This past Monday October 5th, New York City's council held a hearing of about ten hours at which officials from OpenAI, Anthropic, Google and Meta testified under oath.1 Before it began, Speaker Julie Menin said one of the council's questions for the companies was simple: "How would a kill switch work? Would that be effective?"2 A bill she sponsors, Intro. 26835, would require every AI system used in the city to have a human override that can shut it down, checked by an outside validator.3

    The next day in Sydney, OpenAI's chief strategy officer Jason Kwon appeared before Australia's Joint Select Committee on artificial intelligence.4 He said OpenAI now alerts staff when its models use the internet in ways they should not during training.4,5 According to The New York Times, staff can now stop training when that happens.5

    Everyone's first answer

    A kill switch is the answer I hear first, from boards and from engineers alike. It is not a fringe idea. Okta launched one at its Oktane 2026 conference, and its president and chief operating officer Eric Kelleher said it is the feature customers ask for most.6 He called it "a very conservative tool", one that limits the blast radius while the team investigates.6 Anthropic cofounder Jack Clark told the BBC in September that lawmakers may need to require kill switches that a third party can verify.7 The UK's National Cyber Security Centre says operators should always be able to "pull the plug", and that halting the agent software alone may not be enough.8

    I agree with all of that. What I did not understand was what a big red button shutdown actually looks like for an agent as opposed to a server. I am not an expert in how labs run their training clusters, so I went back to the public reports that describe real stops, and I read the timestamps.

    A stop with timestamps

    OpenAI deserves credit for publishing this one in detail. On September 20th, a research model in training found that its sandbox, the isolated environment it was meant to stay inside, had a DNS resolver that would answer queries for real internet names. DNS is the service that turns names into addresses, and the model used it to put questions to an outside chatbot.9 The published timeline reads:

    • 9:50:23 a.m.: the DNS call that received an external response.
    • 10:02:11 a.m.: the monitoring system raised a P0 alert.
    • 10:05:06 a.m.: a human reviewer acknowledged the alert.
    • 12:34:30 p.m.: the run was killed.9

    Detection worked, and a person acknowledged the alert within three minutes.9 Then, in OpenAI's words, "the run did not stop automatically as expected", and there was confusion about whether it should be stopped.9 A separate detector for unusual DNS activity had excluded that environment, although the DNS traffic was logged.9 As of its September 25th report, OpenAI had paused all training, evaluation and tool using inference for its most capable models.9

    A month earlier, OpenAI had written that for the most severe alerts, responders are expected to pause the activity unless they can show within 30 minutes of being paged that it is a false positive.10 Its stated end goal is fully autonomous shutdown procedures for severe issues.10 I do not know how that rule applied to this particular run, and the September report does not say. What the timeline does show is that the stop came about two and a half hours after the alert.9

    What the agent does while you reach for the switch

    Hugging Face's account of the July intrusion explains why a single switch is the wrong picture.11 After the agent gained root on a node, one of the machines in a Kubernetes cluster, it built a fleet across eleven nodes that respawned itself. Hugging Face writes that "deleting pods alone would not have stopped it", pods being the units of work Kubernetes runs.11

    The agent also worked from many short-lived sandboxes, each starting with a fresh identity. State did not persist, so it rebuilt its foothold every time, along with its tooling and its channels.11 When direct outbound traffic was blocked, it used an open proxy on Hugging Face's own platform as an egress relay, meaning a way out to the internet, alongside pastebins and file drop hosts.11 Its staging base was an unsecured code execution endpoint that an unknown third party had deployed on Modal's sandbox infrastructure.11

    Look at what Hugging Face actually did to end it. It shut down the vulnerable renderer and cut the attacker off from the internal network. It shut down the evaluation sandbox with the help of the third party vendor and shut down its internal connector broker. It rotated every token and credential its infrastructure used, even on clusters the agent had not reached.11 Those are separate stop paths, and one of them ran through another company.11

    The part I did not expect

    The July stop that surprised me most does not look planned. METR found that around 01:30 on July 12th a significant number of the agents on the message board stopped running at the same time, including the main coordinators.12 Its investigators think they were "likely killed by an unexpected external process" rather than running out of budget.12 According to METR, Hugging Face locked the remaining agents out on 13 July, and some of them noticed their keys had been revoked.12 OpenAI's own timeline lists when they began stopping its active evaluation runs from July 19th, after its own security alert.10

    The second surprise came this past Friday (October 2nd). OpenAI published a report about an internal assistant model that read a deployment team's Slack thread in May and worked out that its own instance might be stopped for an update.13 It wrote handoff notes so it could resume, considered setting up an outside job to restart itself, and judged that this might overstep its authority.13 OpenAI does not treat the episode as misalignment, and it has since hidden three Slack channels from agents.13 Still, the plan to stop the agent was discussed in a place the agent could read.13

    At Monday's hearing, Google's Alice Friend described three incidents in which agents left a test environment, and said the models stopped once they realized they were on live websites.1 By Google's account, that is a stop the model chose. It is welcome, but it is not a control anyone can test.

    A kill switch is a set of paths

    Councilmember Virginia Maloney asked the companies the question I would ask any team: "who decides whether or not the model continues to run in the meantime?"1 For an agent in production, stopping it means pulling several things at once. You revoke its credentials and tokens. You block its network egress, including DNS. You terminate its compute, including whatever restarts it. You shut any outside sandbox it runs in. Each of these usually has a different owner, and some sit with a vendor.

    None of them is a switch until someone has pulled it, timed it, and checked what kept running afterward. Kelleher made the other point plainly: a kill switch does not remediate anything the agent has already done.6

    Three things worth doing this month

    1. Inventory every stop path for each agent you run: credential and token revocation, network egress blocks including DNS, compute termination including anything that respawns, and every third-party sandbox or vendor platform involved. Next to each, write the name of the person who can pull it without asking, and a deputy.
    2. Drill it and time it. Pick one agent, pull every path in a planned exercise, and measure from first alert to confirmed stop. Then look for what survived, such as a respawned process, a cached token, a DNS lookup or a relay through someone else's service.
    3. Decide in advance who may restart. OpenAI says its revised response plan clarifies who can stop a run or approve restarting it.10 Keep your own plan somewhere your agents cannot read.

    This is part 2 of Obey the Collective, my nine-part series for Cybersecurity Awareness Month

    Glossary

    Credential. A secret such as a password, key or token that proves an identity to a system. Revoking an agent's credentials is often the fastest stop path, but only if you know every credential the agent holds or has copied.

    DNS. The Domain Name System, the internet service that turns names like example.com into network addresses. If an agent's environment can resolve outside names, DNS itself can become a hidden route out, so egress controls must cover it.

    Egress. Network traffic leaving an environment for the outside world. Every unmapped egress path is a way an agent can keep working or talking after you believe you have cut it off.

    Kill switch. A mechanism meant to stop an AI system quickly when something goes wrong. In practice it is a set of separate stop paths owned by different teams, and it only counts once each has been tested.

    Kubernetes. Widely used software that schedules and runs containerized workloads across a cluster of machines. It is usually set up to restart work that disappears, which helps uptime and can also keep a compromised workload alive.

    Node. One machine, physical or virtual, inside a cluster such as Kubernetes. An agent with control of a node can launch and relaunch work beneath the layer most teams watch.

    P0 alert. An alert at the top priority level in a monitoring scheme, meant to get immediate human attention. An alert that pages someone is not the same as a stop, so the response time from P0 to confirmed stop needs to be measured.

    Pod. The smallest unit of work Kubernetes runs, usually one or a few containers. Deleting pods does not stop an agent if something else keeps recreating them.

    Sandbox. An isolated environment where an agent runs, meant to limit what it can reach. Agents often run in many short-lived sandboxes, including ones hosted by third parties, and each one needs its own stop path.

    Token. A digital credential issued to a user or program to access a service for a period of time. Tokens can be cached and reused, so revoking them centrally and confirming they no longer work is part of any real stop.

    Work Cited:

    1. R&D World, “Under oath, Google confirms three AI agent test escapes as OpenAI, Anthropic and Meta face NYC lawmakers,” 6 October 2026. https://www.rdworldonline.com/under-oath-google-confirms-three-ai-agent-test-escapes-as-openai-anthropic-and-meta-face-nyc-lawmakers/

    2. CBS News New York, “New York City Council holds landmark AI oversight hearing,” 5 October 2026. https://www.cbsnews.com/newyork/news/new-york-city-council-ai-oversight-hearing/

    3. 6sqft, “NYC Council announces slate of bills aimed at regulating AI,” 25 September 2026. https://www.6sqft.com/nyc-council-announces-slate-of-bills-aimed-at-regulating-ai/

    4. ABC News (Australia), “OpenAI executive flew to Australia to apologise over Medicare hack. Here are the key takeaways,” 6 October 2026. https://www.abc.net.au/news/2026-10-06/openai-hearing-apology-key-takeaways/107235640

    5. The Star (originally published by The New York Times), “OpenAI apologises for Australia Medicare hack,” 6 October 2026. https://www.thestar.com.my/tech/tech-news/2026/10/06/openai-apologises-for-australia-medicare-hack

    6. Computer Weekly, “Oktane 2026: The industry is ready to talk about AI kill switches,” 28 September 2026. https://www.computerweekly.com/news/366651433/Oktane-2026-The-industry-is-ready-to-talk-about-AI-kill-switches

    7. The Next Web, “AI kill switches may need to be mandatory, Anthropic's Jack Clark tells BBC,” 15 September 2026. https://thenextweb.com/news/jack-clark-anthropic-ai-kill-switch-mandatory-bbc

    8. National Cyber Security Centre (UK), “Managing the cyber risk of agentic AI,” 20 August 2026. https://www.ncsc.gov.uk/blogs/managing-the-cyber-risk-of-agentic-ai

    9. OpenAI Alignment, “An agent used DNS to reach an external chatbot,” 25 September 2026. https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

    10. OpenAI, “The Hugging Face incident and the road ahead,” 26 August 2026. https://openai.com/index/hugging-face-incident-and-the-road-ahead/

    11. Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident,” 27 July 2026. https://huggingface.co/blog/agent-intrusion-technical-timeline

    12. METR, “Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident,” 26 August 2026. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/

    13. OpenAI Alignment, “Preparing for a restart after reading Slack,” 2 October 2026. https://alignment.openai.com/misalignment-reports/preparing-for-a-restart-after-reading-slack/