Inspiration

AI coding agents are becoming extremely good at turning prompts into working code. But I kept coming back to a different problem: what happens when the agent correctly implements an incomplete prompt?

A developer might ask an agent to "add permanent account deletion," and the agent may successfully delete the account. But the original request may never mention invalidating existing sessions, preventing the deleted profile from being retrieved, or handling personal data stored in external services.

The coding agent did not necessarily fail. The acceptance boundary was incomplete.

That led me to AgentGuard: an acceptance layer between human intent and AI coding agents.

What it does

AgentGuard separates software acceptance into three layers: Requested, Expected, and Observed.

Requested behaviours come directly from the developer's task. AgentGuard can then surface additional acceptance behaviours and ambiguities suggested by the task and repository context. These suggestions are not silently imposed as requirements. The human reviews them and decides what enters the acceptance contract.

Once the implementation is ready, AgentGuard independently verifies supported behaviours against actual observations and produces one of three verdicts: PASS, FAIL, or UNVERIFIED.

PASS means the observed evidence supports the accepted behaviour as represented by the check. FAIL means the evidence contradicts it. UNVERIFIED means AgentGuard cannot reliably establish the behaviour using the supported evidence available to it.

Established failures are then transformed into an evidence-backed Correction Brief for the next coding-agent action.

The broader loop is:

Human intent → reviewed acceptance contract → coding agent → implementation → independent verification → evidence-backed correction → coding agent → re-verification.

The current prototype demonstrates the correction handoff manually rather than automatically invoking an external coding agent.

How we built it

I built AgentGuard with a Python backend and a React/TypeScript frontend.

The architecture deliberately separates authority across different components: the LLM suggests, the human authorizes product intent, the coding agent implements, the execution layer observes, and the deterministic verifier establishes what the evidence supports.

The backend converts accepted behaviours into bounded observation plans and supports mechanisms including HTTP/JSON assertions, test commands, registered checks, composite verification, bounded synthetic inputs, and stateful HTTP observation sequences.

Verification results are collected into an Acceptance Verification Report that preserves the reviewed contract, verdicts, evidence, and uncertainty boundaries.

For established failures, AgentGuard deterministically creates a Correction Brief using evidence already collected rather than asking another model to invent a diagnosis.

The frontend provides a controlled demonstration of the complete acceptance flow, from reviewing suggested behaviours through verification and correction handoff.

Challenges we ran into

The hardest challenge was deciding where AI should and should not have authority.

An LLM is useful for interpreting natural-language requirements and suggesting behaviours a developer may have missed, but allowing the same probabilistic system to silently define requirements and declare them satisfied would weaken the trust model.

Another challenge was handling uncertainty. It would have been easy to treat unsupported behaviours as successful or unsuccessful, but neither conclusion is justified without evidence. This led to making UNVERIFIED a first-class verdict.

Stateful behaviours were another challenge. Requirements such as "existing sessions should stop working after account deletion" cannot always be established with a single request, so I added bounded stateful observation sequences and cross-observation assertions.

Finally, I had to keep the project's claims aligned with what was actually implemented. AgentGuard is not a universal software verifier, and the current prototype does not automatically invoke coding agents or run a fully autonomous correction loop.

Accomplishments that we're proud of

I'm proud that AgentGuard became more than a simple LLM wrapper.

It has a clear trust boundary between probabilistic interpretation, human product decisions, execution, and deterministic verification. It can preserve uncertainty instead of forcing binary answers, produce evidence-backed acceptance reports, and convert only established failures into bounded corrective context.

I'm also proud of building an end-to-end experience that makes the concept tangible. A developer can start with an incomplete request, review potential acceptance gaps, establish a contract, inspect independent verification results and evidence, and prepare the resulting failures for the next coding-agent action.

What we learned

The biggest lesson was that software correctness always depends on another question: correct according to what?

A coding agent can produce technically good code while still missing behaviour that never appeared in the original request.

I also learned that probabilistic reasoning and deterministic verification are useful for different things. Models are powerful for interpreting ambiguous human intent and proposing possibilities. Deterministic systems are better suited to deciding whether specific observed evidence satisfies a represented assertion.

Finally, I learned that sometimes the most trustworthy answer a verification system can give is: "I don't have enough evidence."

That is more useful than manufacturing confidence.

What's next for AgentGuard

The next step is evaluating AgentGuard on a larger collection of real coding-agent tasks and measuring how often it identifies meaningful acceptance gaps, how often developers accept or reject its suggestions, how much of those contracts can be grounded into reliable observations, and whether its correction briefs improve subsequent agent attempts.

I also want to expand the grounding and observation capabilities, strengthen evidence provenance, integrate directly with coding agents, automate correction handoff and independent re-verification, and eventually explore CI integration.

The long-term goal is not to replace coding agents. It is to give them an independent acceptance layer.

The coding agent can act on the evidence. It still doesn't grade itself.

Build with AI. Verify with evidence.

Built With

Share this project:

Updates

Submission history