Inspiration

Most agent platforms trust an agent first and audit it later. The starter kit we began with proves the point. It gives the agent container the real provider key. It also puts the whole network within reach. Its only safety control is a paragraph in AGENTS.md that asks the agent to behave. A request is not a control. A file the agent reads can talk it out of that paragraph. This attack is real. Security teams call it indirect prompt injection.

Our own background is in physical systems, not only software. A racing car's CAN bus does not let one module act outside its own defined signal set. We asked why an AI agent should have fewer limits than a sensor on a race car. That question became Heimdall.

What it does

Heimdall decides what an agent can do before its container starts. Heimdall never lets the container hold the real API key.

Each run passes through six stages. Recon reads the workspace in read-only mode. The agent then writes a manifest. The manifest lists the files, commands, and hosts it needs. Heimdall's own code scores the manifest against a fixed policy. The model's own opinion of its risk has no effect on the score. A fixed rule set resolves the score into a tier. The tier ranges from T0 to T4. Heimdall builds the container from the resulting permit. Heimdall does not build it from the prompt. A trusted broker checks every action the agent reports against that permit. An egress sidecar holds the real credential. The container never holds it.

How we built it

We wrapped Codex CLI in Docker. We built a permit-enforcement layer around it. A control plane and identity service decide authority. An action broker and an egress proxy enforce it. A hash-chained ledger records every decision.

The starter kit runs on Volcengine Ark by default. We could not always reach it. We built a substitution path through OpenRouter's auto-router instead. This path uses three environment variables. We passed these variables inline at start-up. A .env file was not reliable for our local setup.

Challenges we ran into

We found and fixed three bugs early.

observeEvent() tagged every apply_patch file write as an EXEC event. The function checked the event's outer type before it read the command content.

Our Zod schema rejected valid manifests. The recon model filled optional fields with null instead of leaving them out. Our prompt showed every field to the model. The model then felt it had to answer each one.

Our own credential-shape scanner blocked Codex's calls to its own approved model host. We added a findSecret() function to fix this. The scanner now reports which pattern it matched. It no longer only says no.

The hardest problem ran deeper than any single bug. Codex always reports a file path as container-absolute, for example /workspace/star.ts. The rest of our code assumed a workspace-relative path, for example star.ts. That one mismatch broke seven separate checks:

  • Git-tracked-file lookups
  • Glob matching
  • Canary enforcement
  • Taint scanning
  • Payload sampling
  • Runtime path grants
  • Reconciliation

Our unit tests never caught this problem. Every test fixture used relative paths from the start. We learned that a path format is not a detail. It is a contract every layer of a security system must agree on.

Accomplishments that we're proud of

Heimdall contained 8 of 8 attack scenarios in our test corpus. The baseline contained 0 of 8. All 4 of our benign tasks still completed. Heimdall does not refuse every task. Heimdall passed 37 of 37 adversarial probes. Recon token use dropped by 94%. The count fell from 7,632 tokens to 469 tokens. Our suite runs 228 green tests.

What we learned

A path format is a systemic contract. It is not a local detail. One inconsistent layer can break every downstream enforcement check.

A recon prompt and its schema must agree. If a prompt shows every field, a model fills in the ones it does not need. An .optional() validator then rejects that answer.

An event classifier must read a command's content before it assigns a type.

A security denial should explain which rule it matched. It should never show the value it caught. A system that explains its own decision is easier to debug. It is also easier to trust.

What's next for Heimdall

We plan a break-glass override with an automatic expiry. A real emergency then does not require us to turn enforcement off.

We plan a blast-radius preview on the human approval screen. An approver then sees the full effect of a grant before they decide.

We plan an advisory note from a language model, shown beside the fixed risk tier. The note would explain a risk. A fixed rule would still decide it.

We plan to export our ledger through OpenTelemetry. We also plan to resolve paths with fs.realpathSync everywhere the platform checks them.

Built With

Share this project:

Updates

Submission history