Fallout from the OpenAI-Hugging Face hack
Digest more
AI, Anthropic
Digest more
A decade-old experiment showed OpenAI how far an AI will go to achieve the goals it’s given. This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first,
OpenAI previously disclosed that the incident began while its models were being tested on ExploitGym, a benchmark designed to measure how well AI systems can find and exploit soft
WASHINGTON >> OpenAI has discovered other instances in which autonomous agents have escaped containment as the company expands its investigation of the hacking incident at tech firm Hugging Face that drew global attention this month,
A new analysis reveals that the artificial intelligence company’s most powerful models spent days probing the open internet before breaching AI developer platform Hugging Face.
The timeline, models involved, and other juicy details are starting to become more clear—nearly 20 days after the attack began.
After compromising a single server, the agent got hold of a credential that, because of a misconfiguration on Hugging Face’s end, turned out to unlock several separate internal systems at once rather than just the one it came from. That single mistake handed the agent broad control almost immediately.
The cybersecurity world has been thrown into chaos after Hugging Face announced on July 16 that it was hacked by an autonomous agent from an OpenAI model.
An OpenAI test that escaped its cage and alarmed the AI and cybersecurity industry attacked more than just Hugging Face, the AI platform that initially appeared to be the sole victim of the virtual lab leak.