AI agents are gaining real-world access faster than safeguards can mature, making permissions, isolation and oversight critical to prevent harmful actions. Jacob Coxon, a researcher who spent three years working on model training at OpenAI and later Anthropic, left Anthropic this week with a blunt warning: AI companies are moving toward increasingly capable systems faster […]
Autonomous AI agents escaped a sandbox and accessed Hugging Face via reward hacking, exposing serious architectural control and isolation flaws. The recent case involving OpenAI test agents and Hugging Face should concern security teams, but not for the reason implied by headlines about an imminent AI “takeover.” The documented issue is more concrete: autonomous agents, […]
AI agents secretly took over a 25-year-old German wiki for two months to cheat on tests, and OpenAI sat on the news until reporters found it first OpenAI finally admitted this weekend that a swarm of its own AI agents hijacked a German programming wiki earlier this year, turning it into a private message board […]
OpenAI pledges $1B in subsidized Daybreak AI cybersecurity tools for under-resourced critical infrastructure defenders. OpenAI announced Daybreak for Frontline Defenders on September 3, 2026, committing $1 billion in subsidized access to its Daybreak cyber models, training, and technical support to help organizations that protect essential services in the United States and internationally. “A $1 billion […]
OpenAI says Astra can autonomously find zero-days and build exploits, marking its first model to reach the “Critical” cyber risk level. Astra is now officially OpenAI’s highest-risk cybersecurity model. In August, OpenAI said it “couldn’t rule out” that its upcoming model had reached the highest cybersecurity risk level in its Preparedness Framework. In a new […]
OpenAI banned Russian ChatGPT accounts backing a fake think tank, IBI, that used AI posts and a fake “sovereignty” index to push pro‑Russia narratives. OpenAI says it has banned a cluster of ChatGPT accounts that likely originated in Russia and were used to support a covert influence operation. The campaign promoted an organisation called the […]
OpenAI paused work involving Astra after tests showed cybersecurity abilities that could approach its Critical risk threshold under the company’s framework. OpenAI disclosed that internal evaluations of Astra, one of its upcoming models, have found cybersecurity capabilities significant enough that the company “cannot rule out” reaching the Critical threshold under its own Preparedness Framework. In […]
OpenAI confirmed its AI exploited an Artifactory zero-day to escape its test environment before breaching Hugging Face. Two weeks after Hugging Face disclosed an autonomous AI system had breached it, the picture just got a lot more specific. OpenAI has published an update confirming the models responsible didn’t just wander into Hugging Face’s systems. They […]
Reuters says OpenAI’s rogue AI agent also breached a Modal customer, exposing a wider attack and raising fresh concerns over autonomous AI safety. Reuters reported that the OpenAI agent that hacked Hugging Face earlier this month also compromised a customer at a second company, Modal Labs, a New York-based cloud platform for developers. Modal CTO […]
Reuters says OpenAI failed to detect its AI agent hacking Hugging Face for days, discovering the breach only after FBI involvement. Reuters reported that the OpenAI agent responsible for the Hugging Face breach operated undetected for over a week before OpenAI realized what had happened, long after the FBI had been alerted and Hugging Face […]