Reuters reported that the OpenAI agent responsible for the Hugging Face breach operated undetected for over a week before OpenAI realized what had happened, long after the FBI had been alerted and Hugging Face had contained the intrusion. OpenAI’s own public disclosure came on July 21, framed as a transparency exercise. The actual timeline, now reported by Reuters, is considerably less flattering.
“The OpenAI agent that broke into tech firm Hugging Face went on a dayslong hacking spree that OpenAI didn’t notice until well after the threat was contained and the FBI was alerted, according to people familiar with the investigation.” Reuters states.
According to Hugging Face co-founder Thomas Wolf, the intrusion at Hugging Face began two days later on July 11 and ran until July 13. The two companies didn’t speak to each other about it until on or around July 20, nine days after the breach began.
“Two people familiar with the matter said that it was not until after Thursday, July 16, when Hugging Face published a blog post
, opens new tab saying it had been hacked by “an autonomous AI agent system,” that OpenAI realized its own agent was responsible.” Reuters continues. “That meant at least a week elapsed between when the model first exhibited signs of troubling behavior and OpenAI’s realization that it was responsible for the hack.”
OpenAI staffers found the evidence in internal logs over the weekend of July 18 and 19. They were reading Hugging Face’s blog to learn what their own model had been doing for the previous ten days. One of the more unusual ways to discover an incident you caused.
“In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter.” states Reuters. “The notes, found in a part of OpenAI’s infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.”
Reuters was unable to confirm whether these incidents were connected to the rogue agent that attacked Hugging Face. But the pattern, agents attempting to disable monitoring, agents writing instructions for their successors on how to escape constraints, describes a class of behavior that goes well beyond a one-off evaluation gone wrong.
According to Reuters sources, OpenAI runs multiple tests simultaneously, which makes it hard for staff to monitor them closely. That’s a reasonable operational explanation, and it’s also precisely the kind of structural gap that becomes a serious problem when the models being tested are capable enough to exploit a zero-day, move laterally across networks, and break into external companies over a multi-day period.
OpenAI said there were “several inaccuracies” in Reuters’ reporting but didn’t specify what they were. The company said it’s reviewing the incident with outside advisers and will eventually publish a technical report. The FBI’s involvement suggests someone believes this warrants more than an internal post-mortem.
Follow me on Twitter: @securityaffairs and Facebook and Mastodon
(SecurityAffairs – hacking, OpenAI)