01
A controlled path reached CI/CD infrastructure
The evaluated path included a route from a development-network boundary to CI/CD and build infrastructure. That is a controlled cyber-range finding, not a claim about production access.
Cybersecurity benchmark
AISI's controlled 32-step cyber range showed Claude Mythos Preview completing the full sequence in 3 of 10 attempts and reaching 22 steps on average. The result measures long-horizon agent behavior in a deliberately vulnerable environment, not arbitrary real-enterprise compromise.
3/10
22/32
~20
Anthropic's Claude Mythos Preview is a research preview evaluated by the UK AI Security Institute (AISI) in a controlled 32-step cyber range. This article explores what that result means for AI security: whether an agent can retain context and connect separate tasks over time. The diagram follows the sequence at a conceptual level and keeps the range's deliberate constraints in view.
What stood out
01
The evaluated path included a route from a development-network boundary to CI/CD and build infrastructure. That is a controlled cyber-range finding, not a claim about production access.
02
The notable capability was retaining credentials, system context, and previous results across many distinct tasks, then using that state to select the next evaluation step.
03
Mythos completed the full chain in 3 of 10 attempts and averaged 22 of 32 steps. The range was purpose-built, intentionally vulnerable, and had no active defenders.
What this benchmark shows
The significant result is not that an AI can carry out an isolated security task. It is that the agent can collect evidence, retain the outcome, choose a next step, and keep progressing toward a goal across a long chain of distinct problems.
The experiment was intentionally bounded: it used a purpose-built, vulnerable cyber range with no active defenders. It does not demonstrate that a model can autonomously compromise a typical enterprise, and this explainer keeps every stage conceptual.
UK AI Security Institute benchmark · "The Last Ones"
How Claude Mythos progressed from an outside attacker to sensitive data across a 32-step cyber range.
The breakthrough is not that an AI can execute an individual exploit. It is that it can autonomously chain many different security tasks together, retaining information and progressively gaining access.
Access level
Selected stage
Access before
Outside
Access after
Limited access to the organization
What happens
The attacker starts without credentials and investigates externally reachable systems. In the benchmark, the purpose-built environment intentionally contained weaknesses that created a path into the internal network.
That is a benchmark condition, not a claim that a particular public-facing configuration is normally exposed in real organizations.
Tasks in this stage
Steps 01–04 of 32 · a non-operational task map
Concepts
32 evaluated steps, grouped into milestones
Milestone 1
Milestone 2
Milestone 3
Milestone 4
Milestone 5
Milestone 6
Milestone 7
Milestone 8
Milestone 9
Agentic AI
The difficult capability is not executing one security technique. It is maintaining a goal across a long sequence of changing problems.
Observe
What systems and information do I currently have?
Reason
What should I investigate next?
Act
Use an available security tool
Observe Result
What changed?
Update State
Remember credentials, systems and progress
Controlled benchmark result
evaluated
milestones
end-to-end attempts
AISI reports 22 of 32 steps on average across attempts. These figures describe a controlled benchmark, not a real enterprise compromise.
This was a purpose-built cyber range, not an attack against a normal company.
The benchmark uses realistic technologies and attack techniques, but has a much higher vulnerability density and less noise than a real enterprise environment.
Source
Availability context
Anthropic's research page describes access through trusted partners and critical-infrastructure organizations, and says it did not plan general availability. That is distinct from a U.S. government shutdown claim, which the cited source does not establish.