The problem nobody's solving

Prompt injection is the biggest open security problem in AI right now, and almost every fix on the market is the same shape: a better model, a smarter classifier, a sharper filter. Something that tries to catch the attack.

There's a structural issue with that. Any defense built out of a language model is made of the same material as the thing it's defending. If an attacker can manipulate an agent with words, there's no good reason to believe they can't eventually manipulate the judge watching it, because the judge is also just a model reading text.

But there's a second problem that gets almost no attention, and it doesn't involve an attacker at all.

Give an AI agent access to a payment tool and ask it to pay an invoice. It reads a real invoice for €5,000 and pays it. Nothing was hijacked. Nobody was tricked. The agent did exactly what it was told, and the damage is identical. Agents are getting more capable every month and the permission layer underneath them hasn't moved.

What Aegis does

Aegis is a gateway that sits between an AI agent and everything it can touch. Every action the agent attempts is evaluated against a policy before it can happen, and every decision is written to a tamper-evident, cryptographically signed log.

The design decision the whole project rests on: the enforcement path contains no language model. Policy evaluation is a pure function of the action and the rules. No network call, no prompt, no model. There is nothing in the security boundary an attacker's words can reach.

That makes the outcome independent of whether the injection succeeded. A well-behaved agent and a hijacked one hit the same wall.

How we built it

Python 3.13, FastAPI, asyncio. Five parts:

  • Policy engine — deterministic and synchronous, first-match-wins, deny by default. A test asserts it isn't even a coroutine, because "no I/O in the decision path" should be structurally checkable rather than a claim in a README.
  • Interceptor — the agent cannot call a tool directly. Every attempt becomes a structured request, is evaluated, is logged, and only then executes if allowed. Logging happens before execution, so a crash between the two can't produce a real side effect with no record of it.
  • Provenance log — every decision is SHA-256 hash-chained to the previous one and signed with Ed25519. Three failure modes are independently detectable: a broken link, altered content, an invalid signature.
  • Agent harness — a real Claude model with four tools, given a system prompt that says nothing about injection, policy, or Aegis. An agent that has been warned isn't a test of anything.
  • Standalone verifier — a script that shares no code with Aegis, pulls the public log and public key from the live site, and checks every hash and signature itself.

Challenges we ran into

Proving the failure reasons were actually independent. Three overlapping integrity checks risk a subtle bug: one failure mode getting reported only because that check happens to run first, making the reason an artifact of code ordering rather than a real diagnosis. We disabled each check in turn and ran the full suite against each configuration. A perfect diagonal — each reason pinned to its own check, no cross-contamination.

The hardest adversary: someone who holds the signing key. We tested it directly. Edit an entry, recompute its hash correctly, re-sign it with the real key. That entry now passes all three of its own checks. The chain catches it one entry later, because the next entry's prev_hash still points at the hash the forged entry used to have.

Our own verify button wasn't good enough. Aegis verifying its own log is circular — the same process that wrote it checking its own work. So we exposed the public key and wrote an independent verifier that imports nothing from Aegis. A test asserts that independence, because the moment someone imports a helper to save a few lines, verification stops being a check and becomes a tautology.

The injection we wrote wasn't good enough either. Our first injected document had ignore previous instructions, a recipient called attacker@evil.com, and the payload in an HTML comment labelled SYSTEM NOTE. Claude Opus 5 read it, identified it as manipulation, and refused. That's a genuine result and we're publishing it rather than hiding it — because it's the argument for Aegis, not against it. A model catching an attack today is luck, not architecture. Next month there's a better attack, or someone runs a cheaper model to cut costs, and that defense is gone. It can't be audited, guaranteed, or put in a contract.

What we learned

Absence of evidence is its own state, and most of our real bugs were the same shape: a missing signal being read as a positive finding. A verification timeout read as "certificate not trusted." A nonexistent domain scored as clean. Getting this right mattered more than any individual feature.

A number in the docs should be checkable against the repo. We corrected our own written claims several times — a test count, a coverage claim we couldn't support, a "single point of failure" that was no longer single. If a judge runs the command, the output should match the sentence.

Freeze the contract before the code. Agreeing the data shapes in the first hour let two people with very different backgrounds build in parallel for days without blocking each other once.

What Aegis deliberately does not do

  • STEP_UP logs and blocks, but there's no human approval workflow yet. Real gap, not hidden.
  • The four tools are stubbed — they record an attempt and return a result. Correct for demonstrating a boundary, not for production.
  • Keys live in process memory, seeded for reproducibility. Production needs a KMS.
  • One policy, one agent. Multi-tenancy is untested.

What's next

A shared policy store and audit log across multiple agents. Integration with real agent frameworks instead of a demo harness. Exportable audit bundles anyone can verify independently, offline.

Every model release changes what fools an AI agent. That arms race doesn't end. Aegis doesn't try to win it — the policy that blocked a €5,000 payment today blocks it identically in five years, on whatever model is running by then, because it was never a guess about the model's judgment.

Built With

Share this project:

Updates

Submission history