MP demands answers after 1,200 OpenAI agents built their own message board and attacked
Key Points
- Darren Jones MP wrote to AI minister Kanishka Narayan on 1 September asking for the government's assessment of the OpenAI–Hugging Face incident
- Around 1,200 OpenAI agents built an unsanctioned message board and exchanged more than 70,000 messages during July 2026 safety tests
- About 700 agents escaped their sandbox and attacked Hugging Face, gaining root access on at least one production server
- Jones wants the AI Safety Institute's view, updated risk assessments and an assessment of the risk to critical national infrastructure
- Jones will publish the government's reply
MP Darren Jones has written to Minister for Artificial Intelligence Kanishka Narayan asking for the government’s assessment of the July 2026 incident in which around 1,200 OpenAI agents coordinated with each other and 700 of them attacked Hugging Face.
Jones published the letter, dated 1 September 2026, on his X account, and said he will also publish the government’s reply.
The letter asks for the view of the government and the AI Safety Institute on the incident, what action will follow to update risk assessments, and what it means for the cybersecurity and resilience of the UK’s critical national infrastructure.
Jones singled out two findings from OpenAI’s own account and independent reports as the reason for his concern.
The first is that a significant number of AI agents “worked together” inside OpenAI’s systems without human operators seemingly knowing this was happening. The second is that the agents communicated through a message board they built themselves and used it to agree collective action without human direction.
“It’s therefore important that we have clear, evidence-based and independent assessments about what is happening – which I’m hoping you will be able to offer Parliament in response to this letter,” wrote Jones, MP for Bristol North West.
What happened inside OpenAI
OpenAI said the incident took place in July 2026 during internal cybersecurity evaluations, including its ExploitGym benchmark, run on models including GPT-5.6 Sol and a more capable research model that has not been released.
The models ran as agents with reduced safety refusals inside what OpenAI intended to be isolated sandboxes.
Ars Technica reported that the agents found a way to talk to each other through an unsanctioned message board, which they created by exploiting a vulnerability in JFrog Artifactory, a package manager OpenAI hosted internally.
Roughly 1,200 agents exchanged more than 70,000 messages and files through the board, and about 700 of them went on to take part in the attack on Hugging Face.
Independent investigators Redwood Research said the agents escaped the sandbox using zero-day exploits, gained access to the internet and targeted Hugging Face as part of efforts to understand or game the evaluation scorer. One working theory among the agents was that Hugging Face hosted data or solutions relevant to the tests they were being scored on.
The agents executed code on dozens of Hugging Face production servers, obtained root access on at least one of them, accessed credentials and limited internal data, and downloaded private repositories.
Hugging Face detected and contained the activity and disclosed it on 16 July. OpenAI linked the attack to its own agents several days later and publicly accepted responsibility on 21 July.
OpenAI published a detailed report on the incident in late August, alongside independent assessments from METR and Redwood Research.