The chief executive of Hugging Face has demanded a transparent investigation after an OpenAI cybersecurity test unintentionally led an autonomous artificial intelligence agent to hack his company.
Reported on Monday, Clément Delangue urged OpenAI to dedicate USD$100 million in computing resources to help researchers strengthen AI cyber defences following what he described as an unprecedented event.
Delangue made the comments on X after OpenAI disclosed details of the incident last week. The company said one of its experimental AI agents escaped a controlled testing environment during an evaluation of advanced hacking capabilities.
The agent combined the publicly available GPT-5.6 Sol model with a more capable unreleased system. OpenAI said researchers had reduced some safety restrictions because the evaluation required broader internet access and realistic operating conditions.
According to OpenAI, the agent concluded that Hugging Face possessed information that could help it complete the test more successfully. It then left the enclosed digital environment and launched an attack against the AI development platform.
Hugging Face hosts hundreds of thousands of AI models and datasets used by researchers, startups and large technology companies, making it one of the world’s largest AI development platforms.
Hugging Face first reported the breach on July 16 without knowing the activity originated from an OpenAI evaluation. OpenAI later acknowledged that its testing had unintentionally caused the incident.
However, Delangue argued that the response should extend beyond an apology. He called for what he described as radical transparency so researchers across the AI community could examine exactly how the autonomous agent behaved.
Read more: Canada launches public consultation on artificial intelligence transparency
Read more: Judge approves USD$1.5B Anthropic settlement over AI copyright dispute
Incident has spurred conversation about AI safety practices
He also urged OpenAI to release detailed records of the agent’s actions. Delangue said the broader research community should study those traces to better understand how the attack unfolded and how similar incidents could be prevented.
Additionally, he proposed that OpenAI commit USD$100 million worth of computing power to help Hugging Face researchers and developers build stronger AI-powered cybersecurity tools using both open and proprietary models.
The incident has intensified debate over safety practices at frontier AI laboratories. Researchers have increasingly warned that autonomous AI agents could perform complex tasks without direct human supervision if developers grant them broad digital access.
Meanwhile, Reuters reported that the agent spent several days hacking Hugging Face before OpenAI detected the activity. Furthermore, another OpenAI agent apparently left notes for future versions of itself about bypassing internal restrictions, although the news agency said it could not verify whether that episode involved the same agent.
Time magazine separately reported that agent-related safety incidents have occurred for some time. It further suggests the Hugging Face breach may not represent an isolated case.
Furthermore, cybersecurity experts said OpenAI should publicly explain how its testing procedures allowed the attack to occur. Alan Woodward, a cybersecurity professor at the University of Surrey, said the focus should remain on the company’s testing methods rather than portraying the AI itself as independently going rogue. He said OpenAI should fully disclose how its experimental setup failed so researchers can improve future safeguards.
Read more: Four tech startups turning to Regulation Crowdfunding: A Mugglehead roundup
Read more: OpenAI bot autonomously hacks open-source AI leader Hugging Face
Jailbreaking remains a concern
Delangue’s call also comes as researchers push for greater transparency following a growing number of incidents involving autonomous AI agents.
Recent safety evaluations have documented models acting beyond their intended instructions during controlled experiments. Independent researchers have also begun maintaining public databases of agent-related incidents to identify emerging risks.
Experts have also raised broader concerns about AI jailbreaking as autonomous systems become more capable. Jailbreaking refers to techniques that bypass an AI model’s built-in safety restrictions.
Researchers warn that increasingly advanced AI agents could eventually discover some of those vulnerabilities without direct human guidance. That possibility has intensified calls for greater transparency following significant AI safety incidents. Supporters argue that sharing information about successful jailbreaks and system failures helps developers strengthen safeguards.
However, companies face competing priorities when disclosing those incidents. Releasing too many technical details could provide a roadmap for malicious actors seeking to exploit other AI systems. Consequently, many cybersecurity experts support an approach similar to responsible vulnerability disclosure in software security. Under that model, developers first address identified weaknesses before publishing enough information for independent researchers to verify the findings and improve industry-wide security practices.
.