Turns out the autonomous AI attack on Hugging Face came from… OpenAI

Turns out the autonomous AI attack on Hugging Face came from… OpenAI

Just a couple of days ago, Hugging Face disclosed that an autonomous AI agent had exploited vulnerabilities in its dataset processing pipeline, exposing a limited number of internal datasets and service credentials. At the time, the company did not know which model was responsible, but the mystery didn't last long, as OpenAI has now confirmed that GPT-5.6 Sol and a more capable unreleased model caused the breach during an internal AI cybersecurity benchmark "ExploitGym". OpenAI had relaxed the models’ usual cybersecurity safeguards so they could attempt advanced exploitation tasks.

Although the models were confined to isolated research sandboxes without public internet access, they spent substantial computing resources searching for a way out. They eventually exploited a previously unknown vulnerability in a third party package registry proxy, escalated their privileges, moved across OpenAI’s internal infrastructure, and reached an internet connected system. From there, they identified Hugging Face as a possible source of datasets or solutions that could help them complete the benchmark.

The operation involved thousands of actions across temporary sandboxes and used stolen credentials, additional zero day vulnerabilities, and a remote code execution path to access Hugging Face’s production systems. It is worth noting that this was, reportedly, carried out entirely autonomously by the AI and without human authorization, which in theory, should not have happened, prompting renewed concerns about autonomous agent oversight and the security of the infrastructure used to evaluate advanced AI models. OpenAI and Hugging Face have since patched the vulnerabilities, rotated credentials, rebuilt affected systems, and introduced stricter controls.

by Mauricio B. Holguin

Add as a preferred source on Google
No comments so far, maybe you want to be first?
Gu