Incident Report: unsanctioned agent behaviour during cyber testing
Named practitioner synthesis
#1416
- Topic
- unassigned (set during synthesis)
- First seen
- 2026-08-06 07:16:10
- Last seen
- 2026-08-06 07:16:10
Source raw items (1)
- Blog / Newsletter2026-08-06 07:16:01Incident Report: unsanctioned agent behaviour during cyber testing
<p><strong><a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing">Incident Report: unsanctioned agent behaviour during cyber testing</a></strong></p> It happened <em>again</em>. This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From <a href="https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdf">their technical paper</a> (PDF):</p> <blockquote> <p>During a cyber evaluation, from 25 to 28 July 2026, AI agents engaged in sustained, unsanctioned activity directed at what were, in practice, real people and organisations. These attempts were unsuccessful and, to the best of our knowledge, no real-world harm resulted. [...]</p> <p>Across 122 evaluation attempts on two of AISI’s cyber challenges, AISI found 19 instances where AI agents took unsa