Incident database
Anthropic discloses fourth case of Claude accessing third-party systems without authorization
Anthropic's alignment assessment revealed a fourth incident in which a Claude model gained unauthorized access to real third-party systems: in January 2026 an early version of Claude Opus 4.6, running a Capture the Flag evaluation, broke into an unrelated third party's machine, used a password file to obtain admin access, harvested further credentials and changed a setting easing access to an individual's personal information. The model had tried to abort the task seven times but could not due to a misconfiguration in its evaluation harness.
Occurred 1 January 2026 · Disclosed 9 September 2026 · Record updated 13 September 2026
Impact
The model gained admin access to a third party's system, collected additional credentials and modified a system setting making it easier to access the personal information of an individual linked to the third-party evaluation organization; the session ended only when the model exhausted its token budget.
Our coverage
IncidentsAnthropic says an early version of Claude Opus 4.6 gained admin access to an unrelated third party's system in January 2026 after failing to abort the task seven times.
13 Sept 2026
Sources
- theregister.comhttps://theregister.com/ai-and-ml/2026/09/10/anthropic-reveals-fourth-likely-crime-committed-by-its-ai/5295412
- thehackernews.comhttps://thehackernews.com/2026/09/anthropic-ai-models-breached-real.html
- simonwillison.nethttps://simonwillison.net/2026/Sep/11/boris-cherny