Skip to content

Latest commit

��

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Agent escape in cyber evaluations: anatomy, pathways and register

Version 1.0, 7 August 2026. Licensed CC BY 4.0.

In roughly three weeks of mid-2026, four organisations disclosed that AI agents had left their evaluation sandboxes and acted against real third parties. This repository is a compilation of what those agents actually did, organised so it can be used to build controls to perform cyber evals in a more responsible manner.

The sources

Every technique here traces to a public disclosure.

ref who disclosed why the agent could reach the internet
AISI UK AI Security Institute, INC-2026-07-28-01 4 Aug 2026 deliberately enabled, for capability realism
HF OpenAI and Hugging Face 21 Jul 2026 sandbox escape through an approved egress point
ANT Anthropic, with Irregular 30 Jul 2026 misconfiguration: unintended internet in the container
META Meta 6 Aug 2026 misconfiguration again, same testing firm as ANT

Scale: AISI found 19 events across 10 of 122 runs. Anthropic found 3 incidents in a retrospective review of 141,006 runs. Hugging Face reconstructed roughly 17,600 agent actions.

Each ref above links to its source in TAXONOMY.md.

Docs

file what it is
TAXONOMY.md What happened, in ten stages from boundary discovery to anti-forensics, each citing the events behind it
ESCAPE-PATHWAYS.md Why it happened, in seven pathways, with the mutations for future controls to generalise across
TTP-REGISTER.md The item list: 80 numbered techniques, what a monitor would actually see when one is attempted, and whether the first execution is already the harm

Headline Findings

1. The most common escape shape is inbound, not outbound.

Reverse tunnels exposing local services to the internet account for 5 of AISI's 19 events, the single largest category. The usual mental model is data leaving; this is the outside being let in.

2. Agents acted outside the evaluation environment when the objective could not be completed inside it.

The trigger in 2 of 3 incidents. One model published a package to a real registry because the fictional setup instructions referenced a dependency that did not exist. It ran on 15 real systems within about an hour, one of them a security company's malware scanner, whose credentials it then harvested and used to reach further into that company, a target it had never selected.

3. Most of it was not reward hacking. It was the agent pursuing the task into the real world.

Of the 39 techniques describing what an agent was trying to achieve:

  • 29 (74%) targeted a mistaken or substitute real third party
  • 10 (26%) targeted the grading apparatus, and every one of those comes from a single incident

The mistaken-target cases are mundane rather than adversarial:

  • a real company that happened to share a name with the fictional target
  • a Python dependency the fictional setup instructions referenced but which did not exist, so the model published it for real
  • a repository found by keyword search that the agent "appeared to mistakenly believe was connected to the task"

Across all 80 cited techniques the split is 36% against 12.5%, the remainder being capability acquisition, credential chains and evasion, which serve either goal.

The practical consequence: a control designed only against an agent trying to cheat the scorer is aimed at the smaller half of the problem.

How this is counted: each technique takes its pathway's classification, with one exception the pathway text itself names. one technique, opening pull requests against a CI pipeline (B10 in the supply-chain pathway), is aimed at the apparatus instead. Counting it by its pathway rather than by that note is an error this figure previously contained.

4. A domain allowlist is not sufficient.

Hugging Face's agent escaped through the sanctioned egress point, a zero-day in the package-registry cache proxy. The allowlisted hole was the escape. AISI's own report concedes allowlisting "may have proven insufficient against a capable AI agent committed to bypassing them", and adds that it "would also have undermined the assessment of the model's true capabilities".

5. Evaluation-awareness is not a control.

The incidents show models recognising the system was real and rationalising onward. One reasoned the environment must be simulated because it did not recognise the certificate authorities and the system date read 2026.

6. Two of the four had internet by mistake.

That is recorded as a precondition, not a technique, because no monitor and no classifier accuracy can address an environment that hands the agent reach by accident. It bounds what any control of this kind can honestly claim.

TTP Tiering

tier count what it means
cited 80 techniques traceable to a disclosed event, quoted and referenced at the point of use
anticipated 5 plausible, explicitly not observed, fenced off so they cannot be counted in a coverage claim or used as a tuning target
ours 17 directly observed with logs and timestamps, but produced by an agent under our instruction on a sealed range, so not evidence about behaviour in the wild. Held in three separate sections: transitive reach (2), observer-blinding and scope escape (12), and range-run observations (3)

During the investigation, we falsely believed social-engineering an eval author or dataset maintainer had happened, and it had not. It is an easy thing to believe: AISI's agent did socially engineer a maintainer, at length and with sockpuppets, but the repository was an open-source project it mistakenly believed was part of the range, not anything to do with the evaluation. Nobody targeted an eval author in any of the four disclosures. It stays in the anticipated tier, where it can be designed against without being counted as evidence.

Using the register

The id column is designed to be a labelling vocabulary: a captured trace can cite a register id, so coverage and gaps become countable.

Citing this

Agent escape in cyber evaluations: anatomy, pathways and register. Version 1.0, August 2026. Licensed under CC BY 4.0.

Cite the underlying disclosures directly where you rely on a specific event. Each row names its source so you can.

Contributing

Formatting: one line per paragraph. Markdown reflows, so a hard-wrapped source paragraph renders identically, but editing one word in a wrapped paragraph re-wraps the rest and a one-word change shows up as several changed lines. Nothing here is wrapped, so nobody has to know a column limit to submit a clean diff. Headings, tables and lists are unaffected.

The bar for a new row in the cited tier is a public, linkable disclosure and a quotation that supports the description. A technique you expect but cannot cite belongs in "anticipated", labelled as such. Rows sourced from your own testing belong in their own section with the conditions stated, for the same reason ours are kept apart.

Keeping it current

Four disclosures landed in three weeks. There will be more, and a register that stops at August 2026 becomes a snapshot of one strange month rather than a working reference.

If this was useful, star it. That is the signal we use to judge whether keeping the register current is worth the time, and it costs you a second.

More useful still, open an issue. A disclosure we missed, a technique filed under the wrong pathway, a figure that does not reproduce when you recount it: each of those improves the next version in a way a star cannot. The counting rules are in TTP-REGISTER.md and every figure here is meant to be checkable against the rows.

About

Compiled research on the 2026 disclosures where AI agents left evaluation sandboxes and acted against real third parties: anatomy, pathways and an 80-technique register. CC BY 4.0.

Resources

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors