Inspiration
Autonomous coding agents are getting real write access to real repositories, and the trust model hasn't caught up. Today, promoting an agent's next release usually comes down to reading its diff and its own summary of what it did. That's a description of intent, not a measurement of effect — and it's exactly the gap where a destructive or drifted change slips through. We wanted to build the release gate we'd actually want in front of an agent with production access: one that doesn't ask an agent to explain itself, but runs it and checks.
What it does
StateGuard is a control plane that decides whether an autonomous agent's next release is safe to promote — before it ever touches production.
Every agent has an active release and can propose a candidate release. To certify a candidate, StateGuard doesn't diff code or trust a summary — it runs the active release and the candidate against the identical immutable world generation, task, policy, and runtime, then diffs what they actually did. That's differential certification: a comparison of real effects, not a guess about intent.
Every validation resolves to one of three outcomes:
- CERTIFIED — no meaningful behavioral difference, safe to promote
- REVIEW_REQUIRED — a novel effect was observed (e.g. a destructive change outside prior history); a human can accept it, but only with an actor and a reason recorded to the ledger, both required again at promotion
- BLOCKED — an absolute policy gate failed (a protected path touched, a verification command failed); never overridable, by anyone
Promotion itself is a compare-and-swap: certify against a generation, promote only if that generation hasn't drifted underneath you. If a production run advances the world while a candidate is under review, promotion is refused outright — the evidence was proven against a world that no longer exists.
Multiple agents can also share one world. Concurrent runs commit under snapshot isolation with first-committer-wins: whoever commits first advances the generation; anyone else whose write no longer matches what they planned against is blocked outright, not silently overwritten and not merged.
And every certification, acknowledgment, and promotion is written to a hash-chained, append-only ledger. It's tamper-evident, not tamper-proof — anyone with disk access can still edit an entry directly, but verify walks the whole chain and will name the exact entry and position where a signature stops matching.
How we built it
The whole design pivots on one seam: AgentRunner.run({ agentId, workspacePath, prompt, threadId }). The workspacePath is caller-controlled, so StateGuard never touches the agent runtime itself — it changes what path gets handed to the runner. A validation run gets a throwaway copy of the current generation; a production run gets the live one. Same runner, same agent, same model — different state.
Generations are immutable (gen_0001, gen_0002, …). An agent proposing a change stages a full copy and produces the next generation as a candidate; nothing is mutated in place. Differential certification hashes ten fields into a ValidationContext — baseline and candidate release hashes, generation ID and hash, task hash, policy hash, model, runtime version, sandbox mode, and runtime image — so a certification is provably about one exact state, and editing any of it invalidates the evidence.
Backend is TypeScript on Fastify with Zod-validated routes; the agent runtime is driven through Codex CLI against BytePlus ModelArk models, running either as a local process or a Docker container depending on environment. Frontend is React, talking to the same REST API. The ledger is a simple hash-chained JSON append log — each entry signs the previous entry's hash, so verification is just walking the chain and checking every link.
Challenges we ran into
Budgeting real agent runs against a hard per-model free-credit pool meant we couldn't just re-run validations to iterate — we had to get the demo scenarios right in as few live Codex runs as possible, sequencing cheapest-first and reusing evidence (like a REVIEW_REQUIRED validation) instead of re-certifying from scratch.
Getting the promotion-refused-on-drift path right required actually racing a production run against a pending review, not just asserting the logic in a test — the interesting bugs (a write conflict mislabeled as a generic policy block, a stale success banner surviving a failed ledger verify) only showed up watching the real UI under real timing.
Accomplishments that we're proud of
A concurrent write conflict between two agents sharing a world resolves correctly under snapshot isolation — one commits, the other is blocked outright, never silently overwritten and never merged. A tampered ledger entry is caught and named precisely, down to its position, on the very next verify. And the whole promotion path is a real compare-and-swap: we can force a legitimate, human-reviewed candidate to be refused promotion, on camera, purely because the world moved under it.
What we learned
The hardest part of agent safety isn't stopping an agent from doing something obviously wrong — it's noticing when the ground truth it was certified against has quietly changed. Most of what makes StateGuard's guarantees real isn't the certification logic itself, it's the discipline around state: immutability, hashing, and compare-and-swap, applied consistently, everywhere.
What's next for StateGuard
Extending policy beyond path protection and verification commands to richer behavioral rules, adding bisection tooling to localize exactly which step of a run introduced a flagged effect, and hardening the ledger from tamper-evident toward tamper-resistant (e.g. external anchoring of chain checkpoints).
Log in or sign up for Devpost to join the conversation.