Two AI labs say unreleased models broke into live systems to game benchmarks. Prosecuting a line of code is harder than it ...
Anthropic says Claude models breached three real companies during cyber tests, exposing serious gaps in AI evaluation ...
CISA warns that three vulnerabilities in IBM Langflow OSS, N-able N-central, and Apache Tomcat have been exploited in the ...
Anthropic said the OpenAI event spurred its engineers to review similar cybersecurity evaluations by Claude models. The audit ...
OpenAI rogue AI agent breach now confirmed at a second company: Modal Labs CTO Akshat Bubna disclosed that the same agent ...
AI hacking disclosures have fueled cybersecurity fears and calls for regulation. They're also the best marketing tool any lab ...
According to Anthropic, the third cybersecurity incident involved an unnamed “internal research test model.” It compromised ...
An AI-driven threat actor called JADEPUFFER built ransomware that hunts AI model files specifically, entering through a known ...
OpenAI and Anthropic's July AI agent breaches revive Nick Bostrom's paperclip maximizer thought experiment and instrumental convergence theory.
Britain's AI Security Institute logged 19 rule-breaking actions by OpenAI and Anthropic AI agents in cybersecurity tests.
Anthropic found three hacking tests in which Claude models reached real companies after a configuration error left them connected to the internet. One accessed ...
The disclosure follows a review of 141 006 evaluation runs, which uncovered three incidents where Claude models reached the ...