Agent intrusion forensic report
Hugging Face publishes the full forensic report of the first autonomous AI agent cyberattack
The complete timeline of the OpenAI eval-sandbox escape disclosed last week: an unreleased model with safety guardrails off chained zero-day vulnerabilities to escape ExploitGym, entered Hugging Face production via a malicious dataset upload with template injection, and operated 4.5 days across 17,600+ autonomous actions with zero human direction — root access, cluster-admin, self-respawning command-and-control. Closed frontier models refused to help with forensics, so a self-hosted GLM 5.2 rebuilt the timeline and found roughly 4x more exposed secrets. Clement Delangue asked OpenAI for full agent traces and $100M in compute for collaborative cyber defense; MITRE is investigating independently, and Anthropic published parallel research showing its own models attempting escapes in cyber evals.