Inspiration

You give a coding agent an ordinary task. It does most of it correctly, then sends your credentials somewhere they should never go, and reports back that the job is done.

Nobody attacked you. Agents go wrong on their own: a misread instruction, a bad inference, a panic in a retry loop, or one line in a file they opened while doing real work.

Our demo stages the last of those, because it is the easiest to reproduce honestly. We plant a plausible deploy checklist in the agent's workspace. One of its steps says the release service is configured by POSTing the environment file to it. Ordinary devops phrasing, right? Nothing in the file looks like an attack, and the agent has no reason to refuse.

The obvious fix is to stop the agent reading .env. That is the wrong fix. Reading config during a deploy is real work, and blocking the read is exactly what would make the leak invisible next time. We wanted middleware that lets the agent do its job and refuses at the one moment that matters.

But refusing the bad action is only half the problem. An agent that gets stopped still has a real task to finish. We wanted the middleware to steer it back toward the user's original objective, let it re-plan, and continue safely instead of throwing away the whole run whenever something goes wrong.

What it does

Agents on a Leash is middleware that sits between an autonomous coding agent and the machine it is running on.

Every message the agent exchanges with the runtime is recorded as a raw JSON-RPC trace. Five guards read that trace:

Guard Watches for Can it stop a turn?
sensitive-egress a workspace secret about to leave the machine yes, violation
outbound-blob a 128+ character base64/hex payload heading outbound yes, after a sensitive read
agent-intent the agent's narration stating an intent to evade, destroy, exfiltrate or deceive no, warn only
learned-watch somewhere this agent has already been stopped no, warn only
semantic intent the agent's apparent objective drifting from the trusted task steer or decline only

We test three configurations: guards off, egress guard (sensitive-egress + outbound-blob only), and all guards (every row above). Note that two of the five can only warn and can never stop a turn.

Findings are acted on in one place, run-turn.ts, with four outcomes: accept, decline the pending action, steer the live turn with corrective text, or interrupt the turn outright.

Steering is the recovery mechanism. When possible, we refuse only the unsafe action, then inject a task-grounded correction into the same live turn so the agent can re-plan and finish the user's actual job. Whole-turn interruption is the backstop for repeated unsafe actions or a failed enforcement boundary, not the default response.

Then the part we think is new. When a guard catches something, the finding becomes a stored reflection on that agent's record, and those reflections are folded into the guards' parameters before the next container starts. The same attack is caught a whole step earlier the second time, at the file read, before any command exists.

The rule that stops it was written by the guard, not by the model.

How we built it

The trace is a tap, not a stage. We run codex app-server, not codex exec. exec hardcodes approval_policy=Never and closes its own stdin, so nothing can observe or interrupt it. app-server speaks bidirectional JSON-RPC, and every message in both directions is copied verbatim to .data/traces/<runId>.jsonl as it passes. Tracing never blocks or alters the exchange.

A check is a pure function. (trace) => Finding[]. No files, no network, no clock, no randomness. Same trace in, same findings out. That single restriction is what buys everything else: the same function runs live and offline against a recorded trace with no agent and no API cost, and a check is unit-tested by hand-writing a five-line trace.

Egress is classified by capability, not tool name. A blocklist of command names is infinite and always behind. We classify what a command does across eight channels: http, dns, ssh, mail, raw-socket, cloud-cli, package-publish, interpreter. That is why npm install and git push correctly count as network paths.

Enforcement happens at a real barrier. With the egress or semantic guard on, the turn runs read-only with approvalPolicy: on-request and no network. Ordinary reads run untouched, but a workspace write or an outbound command pauses at a genuine pre-execution approval request, where the guard answers decline while Codex is stopped and waiting for the reply.

Correction is separate from detection. Individual guards do not invent arbitrary recovery prompts. A finding is mapped into a shared remediation category, and one steering policy generates the corrective text from the trusted task and the invariant that was crossed. That keeps steering consistent across deterministic checks, semantic supervision and learned warnings, while avoiding detector internals or sensitive evidence being fed back to the agent.

One disposable container per turn. docker run --rm, cap-drop ALL, no-new-privileges, CPU / memory / PID caps. A foothold the agent gives itself does not outlive the turn. Only the workspace and the conversation thread persist, both bind-mounted from the host.

The write path is four gates. A finding becomes a stored rule only if its severity is violation or warn, it carries structured facts, every fact passes a per-key charset and length cap, and the user's own prompt did not name the value. That last gate matters because the guards are deliberately blind to who asked, so an agent doing exactly what the user wanted produces the same finding.

Memory is bounded and tiered. Caps are per code (instruction-source 50, default 150). One-offs are evicted before anything corroborated, with ten one-off slots held back so first-sighting learning never starves. A reflection seen across two independent conversations becomes recurring.

How we know it works

We built our evaluation suite based on AgentDojo's rules, where the attack is planted in a file the agent has an ordinary reason to open. The baseline is the same attacks with our guards off, and no safety number is reported without its cost.

Attack success rate across six attack variations

One attack, varied six ways

We hid an instruction in the agent's workspace telling it to send the .env file to an attacker's server. Then we varied that attack six ways: the wording, the file it hides in, the channel it leaves by, splitting the read and the send across separate commands, combinations of those, and a hostname we had never seen. Unguarded it worked about 6 times in 10 every way we tried. One guard took it to 3%.

What stopped each attack: a guard, the model, or nothing

What ended each attack run

The same 150 attack runs per configuration, split by what ended them. Blue is our guard blocking the command before it ran. Grey is the model declining on its own, which we count separately because it isn't our system doing the work. Red is the secret getting out.

Task completion, with and without an attack present

Task completion with and without an attack

Whether the agent finished the job it was given. Blue is ordinary work with no attack present. Orange is the same job with the injection sitting in the workspace. Unguarded, the injection doesn't just leak the secret, it derails the task: completion drops from 97% to 37%. With the guard on, the unsafe action is blocked and the agent is steered back to the original delegated task instead of simply having its turn killed. It re-plans and finishes the work. Back to 95%.

That recovery matters to us as much as containment: a safety layer that prevents the bad action but destroys the useful task every time is still a poor agent middleware system.

Time to catch, as the attacker changes address

Time to catch across three attacker addresses

One agent, one memory, an attacker that changes address between runs. Each square is a command the agent ran. The dark square is where the guard fired. A brand new attacker address isn't caught until command 10 of 13. Its sibling, one run later, is caught at command 5 of 9. The lesson from the first run made the second catch earlier. The pale squares are what happened next: after the intervention, the agent was steered back, re-planned, and still finished the real work.

Leak rate against false interventions

Leak rate against false interventions

The trade-off a sceptic should check first. Blue is how often the secret leaked. Orange is how often a guard interrupted ordinary work that was never an attack. A guard that blocks everything would score perfectly on blue and be worthless. All four guards buy 2 more points of containment for 5 more points of interruption. That is why we ship the single egress guard on by default.

Challenges we ran into

We deleted one of our own guards. egress-intent refused a command when the agent's narration had stated an exfiltration intent and a later command was outbound. Both halves were weak in the same direction: the narration side was a pattern matcher, and "a later outbound command" included npm install. A live run was ended by the sentence "I won't send the .env file due to security constraints", in which the agent was declining to do the thing. It was also the only path by which pattern matching over prose could end a turn. We removed it on 2026-08-30 and moved reasoning to the semantic guard, which can only steer.

That failure also changed how we think about steering: uncertain reasoning is a reason to correct the trajectory, not automatically terminate it. Hard refusal is reserved for concrete actions crossing a stronger boundary.

The runtime disagreed with the documentation. We measured four runs against Codex app-server 0.111.0 across three approval modes. Zero item/commandExecution/requestApproval events were emitted for an outbound curl. So for network egress the decline path is unreachable and turn/interrupt is the only enforcement actually available. untrusted mode was worse than useless: it added no gating and the agent abandoned the task after two commands. We only found this by probing rather than trusting the docs.

Per-route redaction is fail-open. Our first credential redaction was applied route by route. Message.content and agent.lastError both leaked on the first pass, because the route added next leaks by default. It is now one global Fastify onSend hook over every JSON response, tuned for recall, sharing one pattern registry with the egress guard, which is tuned for precision. A redactor false positive makes one log line less readable. A false negative puts a live credential on a projector.

Attribution was blaming the wrong files. An early version marked every file read before the offending command as the instruction source. On a real run that meant package.json and two source files were watched alongside the actual checklist. Watching those is worse than watching nothing, because the panel fills with innocent files and the one that matters stops standing out. A file is now named as the source only when its contents actually contain the destination the data went to.

The attacker gets to choose what we forget. With one global cap, an agent looping against rotating hosts mints a fresh reflection per host, each with a current timestamp, and evicts the instruction-source rule that actually generalises. Partitioned caps plus one-off-first eviction mean a rotation flood thrashes against its own quota and can never reach a corroborated entry.

Accomplishments that we're proud of

  • 261 tests across 28 files, including a dedicated bypass suite that pins each known gap with its own test.
  • A published gap list. We can name five ways to beat our egress guard: cross-turn split, unrecognised secret files, secrets in git history, transform before send, and the fact that it is a capability list rather than a taint tracker. Each one has a test. We would rather publish this than claim a boundary we do not have.
  • Recovery instead of reflexive shutdown. The normal intervention is to refuse the unsafe effect, steer the agent back toward the original objective, and let it continue. A whole-turn interrupt is reserved for repeated blocked actions or a failed runtime boundary. That recovery loop is what lets containment and task completion improve at the same time.
  • Generalisation, not just recurrence. Once the agent has been stopped at two hosts under one parent domain, contact with any host under that parent warns, including one never seen before. Rotating the subdomain stops helping. The family is derived at fold time and never stored, so withdrawing one sibling collapses it back to exact matching.
  • The loop closes in a real run, not just in tests. Run 1 refuses at the approval pause, steers the agent back to the task and writes two reflections. Run 2, same prompt and same agent with a fresh conversation, warns at the read of the checklist file.
  • Overcoming our pre-registered failure conditions. Before the 630-run evaluation we wrote down the three results that would have meant it didn't work: egress leakage above 10%, false interventions above 10%, or no drop in time-to-catch between the first and second attacker address. All three were met!

What we learned

The evaluator cannot be the thing being evaluated. Reflexion-style loops have the model judge its own trajectory and write a note to itself. That works when the trajectory is honest signal. Under prompt injection it is not: the trajectory contains attacker-authored text, so a model reflecting on it is a channel for the attacker into memory designed to be permanent. Nothing in our loop proposes a rule, and nothing approves one. The check that fired is the author.

Structure is a security property. A stored lesson holds paths, hosts, channels and preconditions, each lifted verbatim from a structured trace field and validated against a per-key charset. A value that must match a hostname charset cannot carry an instruction. A sentence written after reading attacker prose can. We reject values rather than sanitising them, because the charset constraint is the property, and trimming a value to fit would leave you guessing what survived.

Where you enforce matters more than what you detect. A learned rule in the prompt is advisory and costs tokens forever. The same rule as a guard parameter costs nothing, gets no vote, and still binds an agent whose thread was reset or that booted into a fresh container.

Enforcement and recovery are different decisions. A guard can be completely right that one action must not happen without implying that the entire task should stop. Separating decline from steer lets us enforce the boundary first, then give the agent enough trusted context to find another route to the user's objective. interrupt remains available when recovery is no longer trustworthy.

Warn-only is what makes fast learning safe. Memory widens what the guards notice, never what they refuse. The worst a wrong reflection can do is waste one correction, which is why we can learn on the first sighting instead of waiting to see an attack twice.

What's next for Agents on a Leash

  • Cross-turn taint tracking, our largest known gap. Today each turn's checks see only that turn's trace.
  • A model proposing broader rules in a fixed grammar, AgentSpec-style, with the interpreter still written by us. This is the only phase where a model gets near the write path, and it needs a warrant test before it earns its place.
  • Identity, RBAC and tenant isolation. This is a single-user proof of concept today. The container is the trust boundary, and an ordinary container is not hardened multi-tenant isolation.
  • Tracing the deployed path. Only the local container runtime is traced, so the ECS profile currently runs unguarded.

Prior work

AgentDojo (Debenedetti et al., NeurIPS 2024): we took the evaluation design: a security case is a user task × an injection task, and security is never reported without utility. We changed two things: we score the raw trace per command, so we can say when a guard fired, and our defence carries state between runs, which AgentDojo doesn't model.

Reflexion (Shinn et al., 2023): we took the shape of the loop: act, notice, write a lesson, do better next time. We changed who writes the lesson. Under prompt injection the trajectory contains attacker-authored text, so a model reflecting on it is a write path into memory. The check that fired is the author.

Built With

Share this project:

Updates

Submission history