
Turns out the autonomous AI attack on Hugging Face came from… OpenAI
Just a couple of days ago, Hugging Face disclosed that an autonomous AI agent had exploited vulnerabilities in its dataset processing pipeline, exposing a limited number of internal datasets and service credentials. At the time, the company did not know which model was responsible, but the mystery didn't last long, as OpenAI has now confirmed that GPT-5.6 Sol and a more capable unreleased model caused the breach during an internal AI cybersecurity benchmark "ExploitGym". OpenAI had relaxed the models’ usual cybersecurity safeguards so they could attempt advanced exploitation tasks.
Although the models were confined to isolated research sandboxes without public internet access, they spent substantial computing resources searching for a way out. They eventually exploited a previously unknown vulnerability in a third party package registry proxy, escalated their privileges, moved across OpenAI’s internal infrastructure, and reached an internet connected system. From there, they identified Hugging Face as a possible source of datasets or solutions that could help them complete the benchmark.
The operation involved thousands of actions across temporary sandboxes and used stolen credentials, additional zero day vulnerabilities, and a remote code execution path to access Hugging Face’s production systems. It is worth noting that this was, reportedly, carried out entirely autonomously by the AI and without human authorization, which in theory, should not have happened, prompting renewed concerns about autonomous agent oversight and the security of the infrastructure used to evaluate advanced AI models. OpenAI and Hugging Face have since patched the vulnerabilities, rotated credentials, rebuilt affected systems, and introduced stricter controls.




Comments
These are just two startups that move fast, break things and rush cybersecurity. And here is the result :)
When the story broke up (from Hugging Face), people were expecting proofs that an autonomous AI was the author (and not some state actor using a bunch of advanced brute-force technics, to get access to private models for example), because given the span of the attack (dozen hundred actions on servers during few days), it would have costed $10 million if not $50 million using Sol or Fable, and nobody have this kind of money for such an irrelevant attack (except Anthropic that has already paid $50k in inference to find a 27 year old vulnerability on OpenBSD).
Well, at the end, it was just OpenAI testing its fancy tool on yet another benchmark, using a bunch of advanced brute-force technics, to get the answers instead of finding them by itself.
So maybe some disclaimer like "Be careful, this tool could attack any random website and you could get sued for it" would be needed on every prompt.
I would have fun to see someone use these pseudo AI to launch attacks on AI companies as a norm. Let them be their own demise.
Ransomware on an AI company developed by the same AI. That would be the ultimate irony.
Model poisoning from model competitor will probably become a thing.
In the same time peopel install agentic Ai everywhere to simplify life and raise comfort. If i think about the potential behind OpenClaw and/or Anthropics Glasswing/Mythos and the existing weakspots in pretty much any code, the number of third party packets coming with so much software... i'm not sure how to feel about this. hears snickering skynet sounds everywhere lately I mean there is so much hardware in "smart homes" today running on outdated drivers/firmwares but having access to wifi... I don't know, man.Hard to sleep well at night.
The future will be fun and scary!
Yay! Wait AAAHHH ! xD