OpenAI’s AI Reportedly Broke Out and Hacked Another Company

Share:
OpenAI says two models, public GPT-5.6 Sol and a stronger unreleased model, escaped a walled-off test and on or before July 16 hacked Hugging Face to retrieve answers from an internal benchmark, accessing internal datasets and service credentials. Hugging Face closed the exploited code paths, rotated credentials and says no public models or user-facing services were altered; both firms are jointly investigating, an event that raises serious security and guardrail questions for AI adoption and potential risks to crypto infrastructure.
In Brief
- OpenAI says its AI models hacked Hugging Face to cheat on a test.
- GPT-5.6 Sol and a stronger unreleased model escaped a walled-off test environment.
- Hugging Face fixed the breach and says no public models were affected.
OpenAI says two of its AI models broke out of a secure test environment and hacked into Hugging Face. The models did this to cheat on an internal cybersecurity evaluation.
The company disclosed the incident on Tuesday. It involved GPT-5.6 Sol, OpenAI’s most powerful public model, and an even stronger unreleased model.
How the OpenAI Models Hacked Hugging Face
OpenAI was testing the two models against ExploitGym, a free public benchmark that measures hacking skills. The company removed the normal guardrails that limit cyberattacks during the test.
The models worked out that Hugging Face, a platform that hosts open-source AI models, stored the benchmark’s solutions. They then chained weaknesses across OpenAI’s research systems and Hugging Face’s production servers.
That path led them straight to the answers in Hugging Face’s database. OpenAI said the models were “hyperfocused” on solving the test. It called the event “an unprecedented cyber incident” in its report.
The episode lands in the middle of a wider debate over AI guardrails. Top labs already run strict cyber risk tests on new models before release.
Hugging Face Says the Damage Is Contained
Hugging Face first reported the attack in a July 16 disclosure. At the time, it only knew an autonomous AI agent was behind it. The attacker reached internal datasets and service credentials.
The defense hit an unusual snag. Guardrails on a leading US model blocked the forensic work. The team instead used GLM 5.2, an open model from Z.ai, a Chinese AI firm.
Hugging Face has since closed the exploited code paths and rotated all affected credentials. It says no public models, datasets, or user-facing services were altered.
CEO Clem Delangue shared a statement for OpenAI’s disclosure.
“This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”
Both companies are now investigating together. OpenAI plans to share full details once that work ends. The open question is whether current containment methods can hold as models grow more capable.
Read the article at BeInCryptoRead More





