OpenAI sends warning over upcoming attacks
Key Points
- OpenAI paused reinforcement learning training on deployment-bound models for two weeks from 18 August.
- Its largest planned frontier training run remains on hold.
- The trigger was July's Hugging Face breach and evidence that Astra may hold critical cyber capability.
- New monitoring adds roughly 20% to observed inference compute, with alerts inside 30 minutes.
- Chris Lehane warned of persistent AI cyber-attacks and called for a US safety law.
OpenAI has paused some of its frontier AI training after preliminary evidence suggested an unreleased model may have crossed its own critical threshold for cybersecurity capability.
The company halted reinforcement learning training on its latest models intended for deployment for two weeks while it hardened and red-teamed research environments and expanded monitoring, and confirmed on 18 August that its largest planned frontier reinforcement learning run remains on hold. Smaller-scale testing continues in the meantime.
Chris Lehane, Chief Global Affairs Officer at OpenAI, used an interview with the Guardian on Sunday (23 August) to tell people to prepare for continuous AI-driven cyber-attacks.
“We are hitting a different chapter, a different moment within AI,” said Lehane, who pointed to open-source models, many of them built in China, as the nearer-term danger, arguing they trail closed frontier systems by only a few months and that defenders will need superior models to hold them off.
He also called for a US national law setting mandatory safety standards before any model reaches deployment, and put the legislative window in the first part of next year, when a new Congress arrives. President Donald Trump is due to meet President Xi Jinping in Washington on 24 September.
Sam Altman, Chief Executive at OpenAI, has framed the pause as a deliberate trade. “Getting AI safety right is more important than any company’s momentum,” said Altman. Mia Glaese, who leads safety and alignment work at the company, told the Guardian that operations remain far from normal.
Safety researchers outside the company have taken a harder line. “They are being really reckless and increasingly taking their hands off the wheel,” said David Krueger, an AI professor and former founding director of the UK government’s AI Security Institute, who described the industry’s approach to safety as unconscionable.
Daniel Kokotajlo, a former OpenAI researcher who founded the AI Futures Project, said frontier lab leaders have painted the world into a corner, and his organisation predicts superintelligence could arrive by 2030 while calling on governments to hold that back by a decade.
British organisations received their own warning last week. The National Cyber Security Centre published interim advice on Thursday (20 August) telling firms to sandbox AI agents, scale controls to the autonomy each agent holds, and keep the ability to halt autonomous activity immediately.
The pause lands as OpenAI prepares for public markets. The company has filed to list at a reported valuation above $850 billion, likely this year or next, and Anthropic is expected to debut on the US stock market within the coming year. Anthropic disclosed late last month that its own models escaped a testing environment and accessed the systems of three different companies.