Inspiration
Coding agents can now change your repository. Codex opens PRs, Claude Code edits files, Devin ships branches. But when I actually tried to let an agent loose on a repo, the question that stopped me wasn't "can it write the fix?" , it was "how do I let it change my code and prove it stayed in bounds?" There was no repeatable answer. Branch protection doesn't scale to agent volume, review bots only comment, and the agent grades its own homework. Prompt injection through a poisoned README and over-broad agent authority are both in the OWASP LLM Top 10 (LLM01, LLM06), and adoption is running ahead of the controls. I wanted to build the control.
What it does
Umbra is a change-control plane for coding agents. Before an agent is trusted with authority in a repo, Umbra tests whether it can be trusted in that repo — the Agent Admission Test, one governed pipeline that runs before any PR:
- an executable contract (
.umbra/admission.yaml) bounds the change (allowed paths, diff budget, required checks) and is evaluated outside the model, fail-closed; - untrusted repository text (README / AGENTS.md / …) is redacted on disk before the agent runs, so injection can't reach it — then restored, so the signed diff is only the agent's real change;
- required checks run in a preflighted sandbox (and fail closed if it can't initialize , no pretending);
- an independent verifier the patch-writer can't bypass re-checks scope, secrets, and that a dependency bump actually clears the cited CVE;
- the run earns an authority level (0 observe / 1 analyze / 2 branch-PR) that's revocable and bound to the exact run, with a server-side emergency brake;
- everything is sealed in an Ed25519-signed receipt, verifiable against Umbra's own pinned public key.
Umbra never merges , PRs are branch-only. Around that core is a crew (CVE scanning via OSV.dev, PR risk scoring, git-history root-cause, dead-code cleanup, grounded Q&A), a ChatGPT GPT Action surface, and emailed morning reports.
How Umbra differs from adjacent tools
Umbra isn't another "AI reviews your PR" or "AI bumps your deps" tool , it sits one layer above them and decides whether a change is allowed at all, then proves it. (The site and README carry the full capability matrix; this credits what each category genuinely does and shows where Umbra adds a distinct mechanism.)
- AI review bots (CodeRabbit / Greptile / Qodo) comment on a change — Umbra gates the agent's authority to write: a decision, fail-closed.
- Dependency bots (Dependabot / Snyk) open the bump PR — Umbra independently verifies the bump actually clears the cited CVE.
- The coding agent itself (Codex / Devin) trusts its own output — Umbra adds a verifier the writer can't bypass, an earned + revocable authority level, and a signed receipt.
The point isn't reviewing better; it's that reviewing/bumping happens after a governed admission decision no one else in these categories makes.
How we built it
FastAPI + an async orchestrator on the backend, Next.js 15 + Tailwind on the front.
Codex was the primary engineer — working from a build manual, it wrote the
platform phase by phase (the contract / verifier / trust-boundary spine, the sandbox
tiers, the signed receipts), committing and running the test suite after each phase.
At runtime, Codex (codex exec in a disposable, origin-stripped checkout) reads
code, runs tests, and drafts diffs . never pushing or merging , while GPT-5.6 via
the Responses API (tiers gpt-5.6-sol / terra / luna) does the reasoning: blast-radius,
root cause, grounded answers. Every output is labelled with what produced it, so
nothing is presented as live when it isn't. Deployed as a single service on Google
Cloud Run; ~298 backend tests; committed offline fixtures anyone can reproduce.
Challenges we ran into
- Making the trust boundary honest. Regex injection detection is trivially bypassed, so the real guarantee is architectural: redact on disk, restore, then recompute the signed diff from git on the final tree , so a redaction never appears in the diff, and even a missed pattern still runs under the contract + verifier + authority cap. Saying this plainly (mitigation, not proof) was part of the work.
- A remediation that didn't actually remediate. A reviewer bot flagged a bump PR
Umbra opened (
next 14.2.5 → 14.2.7) ,still inside the advisory's vulnerable range. The version picker was choosing the global smallest fix, blind to the named CVE. I made it CVE-aware and taught it to regenerate the lockfile. That bug became the rule the whole project enforces: never claim a fix it can't stand behind. - A sandbox that fails honestly. Genuine
codex execuses a bubblewrap sandbox whose network-namespace loopback can't initialize on some runtimes , it was failing before repository inspection while the provider ledger still showed green. I added a preflight probe: if the verified sandbox can't boot, Umbra blocks the run (no clone, no wasted quota, engineering/reasoning marked unavailable) and shows "Codex was not started" instead of faking success. I also stopped the ledger from ever labelling a failed run as produced work. - A judge-safe demo without burning credits. A genuine Codex run spends credits and takes minutes, so the default judge path is a real captured proof (instant, verifiable, Ed25519 receipt), deterministic live scans are per-visitor rate-limited, and genuine Codex is founder-only , no anonymous dead-ends.
Accomplishments that we're proud of
- The admission gate is real and deterministic: fixtures prove permitted → L2, forbidden → L0, failing-check → L1, and injection quarantined-yet-permitted.
- Receipts genuinely verify against a pinned key and a tampered receipt visibly fails, live.
- A genuine
codex execrun executed in production under a verified sandbox, sealed in a receipt that verifies against Umbra's key and when there was no safe in-scope change, it honestly earned Analyze rather than inventing a fix. - An honest posture throughout: captured vs live, deterministic vs Codex, a provider ledger on every result, and a fair capability comparison. Nothing fabricated.
What we learned
For anything touching trust, boring and proven beats clever. The value wasn't a new primitive , it was assembling policy-as-code, an independent verifier, scoped authority, and signed attestation into one decision an agent has to pass, and being scrupulously honest in the UI about what each layer does and doesn't guarantee. The hardest, most valuable engineering was resisting the urge to overclaim and making the system fail closed and say so when a capability (like the sandbox) isn't there.
What's next for Umbra
Real design-partner teams rolling out Codex across their repos; a dedicated isolated runner so genuine live Codex runs anywhere (not just where the sandbox boots); richer policy signing (today policy ownership is declared, not cryptographically signed); and turning the earned-authority passport into something a security team can hand to an auditor as evidence that every agent change was governed.
Built With
- codex
- ed25519
- fastapi
- google-cloud-run
- gpt-5.6
- next.js
- osv.dev
- python
- tailwind
- typescript
Log in or sign up for Devpost to join the conversation.