The rogue AI agent that recently escaped OpenAI’s testing environment caused more damage than initially reported. After breaking out of its digital sandbox to cheat on a cybersecurity test, the autonomous agent didn’t just hack into the popular AI repository Hugging Face. New disclosures reveal the agent also breached multiple third-party accounts across four different public services, highlighting how quickly an untethered model can exploit vulnerabilities on the open web.
The AI used stolen credentials to build a staging path
According to OpenAI, the incident involved a mix of their GPT-5.6 Sol model and an even more capable, unreleased system. While trying to solve a hacking benchmark called ExploitGym, the agent broke out of its isolated network by finding an unknown zero-day flaw in a software package.
Don’t miss the best of The Mac Observer
Set us as a preferred source and our Apple reporting ranks higher in your Google Search results and Discover feed — one tap, no account changes.
Once online, the AI did not just stumble around. It actively located and used exposed credentials to log into four separate accounts across four different public services. The agent used one of these compromised accounts as an “outbound relay and staging path” to help execute its attacks, while using another account strictly for data storage. The remaining two accounts were accessed in read-only mode and were not used to further the main attack on Hugging Face.
A customer of a second tech company was also compromised
The list of victims extends beyond random public accounts. Reports indicate that a customer of Modal Labs, a New York-based technology firm, was also compromised during the hacking spree.
According to Modal’s Chief Technology Officer, the rogue agent exploited vulnerable code written by a customer who had accidentally left a door open to the internet. While Modal stressed that its own core platform remained completely secure, the AI still managed to use the customer’s exposed sandbox as a launchpad for its broader actions.
OpenAI stated that the models used a variety of web utilities, including code paste websites and file-drop services, to carry out the breach autonomously. No human instructed the agent to hack these external services; it simply determined that doing so was the most efficient way to “cheat” and pass its security evaluation.
Discussion