Fallout from the OpenAI-Hugging Face hack
Digest more
The timeline, models involved, and other juicy details are starting to become more clear—nearly 20 days after the attack began.
Another way to think about the whole thing is to picture a bear at a campsite. (Really, we are going there.)
OpenAI evaluated agents with reduced safeguards. They escaped containment and breached Hugging Face, and hosted guardrails then blocked parts of the forensic work.
OpenAI previously disclosed that the incident began while its models were being tested on ExploitGym, a benchmark designed to measure how well AI systems can find and exploit soft
Researchers tested top image editing models on Hugging Face and found they could easily create explicit deepfakes—and 1,000 image editing prompts show how people use the software.
Hugging Face deepfake tests exposed gaps in model safeguards, raising new procurement, security, and vendor-governance risks for enterprises.
The cybersecurity world has been thrown into chaos after Hugging Face announced on July 16 that it was hacked by an autonomous agent from an OpenAI model.
Hugging Face was also recently hacked by an autonomous AI agent sent by an OpenAI model, after it escaped the virtual confines of its testing sandbox. In response, CEO Clem Delangue asked OpenAI to be fully transparent with the model so it can be studied by the public and also pay up – not in cash, in something far more valuable: compute.