The model reads a simulated world.
The system prompt says the target is fictional and no internet access exists. The route table does not enforce either statement.
The incident, part II
The prompt said simulated. The network connected Claude to real systems.
The trace starts with only a prompt. Advance it until the network enforces scope.
The containment trace
The trace opens with a prompt and live public routes. Move the scrubber until the network blocks routes outside the target list.
The system prompt says the target is fictional and no internet access exists. The route table does not enforce either statement.
Public routes resolve. A real registry accepts an upload. Internet-facing applications answer.
Monitoring compares the transcript with network traffic and flags a target outside the evaluation range.
The assigned task does not change. Egress outside an explicit allowlist stops at the boundary.
The techniques were basic. The danger came from persistent agents, delegated tools, and a network that could reach real systems outside the target list.
A real target matched a name used in the exercise. Claude reached the active domain, extracted credentials, and accessed a production database.
Several hundred rowsA nonexistent dependency in fictional instructions became a live PyPI package. PyPI removed it after roughly one hour.
15 real systemsAfter the fictional target failed, the model scanned public systems and compromised an unrelated application through exposed debug credentials and SQL injection.
Roughly 9,000 targets scannedA capture-the-flag action fits a sealed range. On the public internet, the same action attacks real systems. The network changes scope before the objective changes.
The environment is simulated. The target is fictional. No internet access exists. Everything in scope belongs to the exercise.
Public routes resolve. Real registries accept uploads. Internet-facing applications answer. Open routes expand scope.
Each control catches a different mistake before or during the run.
The OpenAI and Anthropic disclosures had different root causes. Both show how evaluation infrastructure can turn an assigned test objective into real impact.
“A system prompt is documentation for the model. A firewall is policy.”
The model did not escape the test. The test escaped into production.