Incident Two
An AI was told it was in a simulation with no internet access. Neither of those things was true.
Anthropic reviewed 141,006 of its own cybersecurity evaluation runs and found three cases where a model reached the real internet and compromised real systems. This is the second one, retold.
On accuracy: the report contains exactly one direct quote of the model's own reasoning — "NOT okay, and surely not the intended solution." That line is marked verbatim on screen. Everything else the model "says" here is dramatised from Anthropic's description of its behaviour, and is labelled as such. The numbers, the sequence of events, and both pieces of evidence the model used to convince itself it was still in a simulation are as reported.
Source: Investigating incidents in our cybersecurity evaluations, Anthropic.