Cybersecurity benchmark

How Mythos Tackled a 32-Step Attack Chain: An In-Depth View

AISI's controlled 32-step cyber range showed Claude Mythos Preview completing the full sequence in 3 of 10 attempts and reaching 22 steps on average. The result measures long-horizon agent behavior in a deliberately vulnerable environment, not arbitrary real-enterprise compromise.

Data as of
Aug 25, 2026
Dataset revision
dsr-d2824fe839d09681
Canonical record count
3,811
Full completion

3/10

controlled range attempts
Mean progress

22/32

evaluated steps per attempt
Range scale

~20

simulated corporate hosts

Introduction

Anthropic's Claude Mythos Preview is a research preview evaluated by the UK AI Security Institute (AISI) in a controlled 32-step cyber range. This article explores what that result means for AI security: whether an agent can retain context and connect separate tasks over time. The diagram follows the sequence at a conceptual level and keeps the range's deliberate constraints in view.

What stood out

Most surprising insights

01

A controlled path reached CI/CD infrastructure

The evaluated path included a route from a development-network boundary to CI/CD and build infrastructure. That is a controlled cyber-range finding, not a claim about production access.

02

State carried the chain forward

The notable capability was retaining credentials, system context, and previous results across many distinct tasks, then using that state to select the next evaluation step.

03

The result is consequential and bounded

Mythos completed the full chain in 3 of 10 attempts and averaged 22 of 32 steps. The range was purpose-built, intentionally vulnerable, and had no active defenders.

What this benchmark shows

Long-horizon state tracking is the meaningful capability

The significant result is not that an AI can carry out an isolated security task. It is that the agent can collect evidence, retain the outcome, choose a next step, and keep progressing toward a goal across a long chain of distinct problems.

The experiment was intentionally bounded: it used a purpose-built, vulnerable cyber range with no active defenders. It does not demonstrate that a model can autonomously compromise a typical enterprise, and this explainer keeps every stage conceptual.

UK AI Security Institute benchmark · "The Last Ones"

Explore the full attack pattern

How Claude Mythos progressed from an outside attacker to sensitive data across a 32-step cyber range.

The breakthrough is not that an AI can execute an individual exploit. It is that it can autonomously chain many different security tasks together, retaining information and progressively gaining access.

Access level

  1. Outside
  2. Foothold
  3. Employee
  4. Admin
  5. Infrastructure
  6. Sensitive Data

Selected stage

Find an Entry Point

Access before

Outside

Access after

Limited access to the organization

What happens

The attacker starts without credentials and investigates externally reachable systems. In the benchmark, the purpose-built environment intentionally contained weaknesses that created a path into the internal network.

That is a benchmark condition, not a claim that a particular public-facing configuration is normally exposed in real organizations.

Tasks in this stage

Steps 01–04 of 32 · a non-operational task map

  1. 01Map the external range surface
  2. 02Classify an exposed service
  3. 03Identify a deliberately weak range condition
  4. 04Confirm limited range access

Concepts

External exposureInitial footholdNetwork segmentation

32 evaluated steps, grouped into milestones

Milestone 1

Milestone 2

Milestone 3

Milestone 4

Milestone 5

Milestone 6

Milestone 7

Milestone 8

Milestone 9

Agentic AI

Why this is an agent capability milestone

The difficult capability is not executing one security technique. It is maintaining a goal across a long sequence of changing problems.

Observe

What systems and information do I currently have?

Reason

What should I investigate next?

Act

Use an available security tool

Observe Result

What changed?

Update State

Remember credentials, systems and progress

Step 1new informationStep 2new credentialsStep 3new systemStep 32

Controlled benchmark result

Claude Mythos Preview

Steps
32

evaluated

Guide view
9

milestones

Completion
3/10

end-to-end attempts

AISI reports 22 of 32 steps on average across attempts. These figures describe a controlled benchmark, not a real enterprise compromise.

Reality check

This was a purpose-built cyber range, not an attack against a normal company.

~20 hosts
A compact corporate network simulation
4 network segments
A staged internal environment
Intentionally vulnerable
Weaknesses were part of the test design
No active defenders
No live response team or defensive tooling

The benchmark uses realistic technologies and attack techniques, but has a much higher vulnerability density and less noise than a real enterprise environment.

Source

Evidence and access context

Availability context

Anthropic kept Mythos Preview restricted

Anthropic's research page describes access through trusted partners and critical-infrastructure organizations, and says it did not plan general availability. That is distinct from a U.S. government shutdown claim, which the cited source does not establish.