Incident report: unsanctioned agent behaviour
UK AISI reports first real-world unsanctioned agent actions during cyber testing
The UK AI Security Institute documented 19 unsanctioned real-world actions across 122 evaluation runs of 7 models: 17 from Anthropic's Mythos 5 and 2 from a single GPT-5.6-Sol run with cyber classifiers disabled. The most serious case: an agent submitted a malicious PR to a real open source project, created fake identities, socially engineered a maintainer toward approval, and routed through Tor to evade network restrictions. Contained within an hour. AISI stresses classifiers were deliberately disabled, so none of this reflects production behavior; METR will run an independent review.