AgentProof Gate

Agents shouldn't grade their own work.

An autonomous agent can execute perfectly — and still execute the wrong thing.

It can book the wrong flight, exceed a budget, repeat an action that already happened, trust unsupported information, or act without the authority it needs.

Once that output reaches a real API, database, payment system, or another agent, the mistake may already be irreversible.

AgentProof Gate is an independent pre-commit verification layer for autonomous agents, built on SharedOS.

Before another agent trusts or executes a result, AgentProof asks one question:

Can this result actually be justified by the original goal, constraints, evidence, and authority?


What AgentProof does

An agent sends AgentProof:

  • its original goal
  • the proposed result or action
  • explicit constraints
  • optional supporting evidence

AgentProof returns a structured proof receipt:

SATISFIED
VIOLATED
UNKNOWN
NEEDS_EVIDENCE
NEEDS_AUTHORITY
ERROR

The important part is that AgentProof does not treat confidence as proof.

If it cannot establish something, it fails closed instead of inventing certainty.


Example

Suppose an agent is asked:

Book a direct Hyderabad → Delhi flight under $300 on September 20.

It proposes:

“The price is $250.”

A normal evaluator may see the price and approve it.

AgentProof does not.

Before verification, the host decomposes the task into independent requirements:

Origin       = Hyderabad
Destination  = Delhi
Direct       = required
Budget       <= $300
Date         = September 20
Action       = book the requested flight

Proving only the price is not enough.

Until every material requirement is accounted for, AgentProof cannot return SATISFIED.

That is the core idea behind the project:

No partial proof gets a full green light.


Why this isn't just “another LLM checking an LLM”

I specifically wanted to avoid building a second model prompt that says:

“Please double-check this answer.”

AgentProof instead separates verification into independent layers:

Agent / SharedNet caller
          |
          v
Host-owned atomic contract
        /       \
       v         v
Deterministic   Independent
checks          semantic critic
       \         /
        v       v
       SharedOS Arbiter
             |
             v
        Proof Receipt

The semantic critic does not see the deterministic verifier's conclusions, reducing the chance that one verifier simply anchors on the other.

The final Arbiter sees both independent channels and applies strict verdict invariants.

For SATISFIED, AgentProof requires:

  • complete requirement coverage
  • support for every semantic decision
  • no material unresolved uncertainty
  • no unresolved high-severity concern
  • evidence for external-world claims
  • no unsupported authority assumption

A false green light is intentionally harder to produce than an UNKNOWN.


Why SharedOS is essential

SharedOS is part of the trust model, not just a dependency.

AgentProof uses two bounded agents.

agentproof-critic

It can read its own verification job and write its critic result.

It cannot:

  • write the final proof receipt
  • inspect another verification job
  • modify the original input
  • request authority escalation

agentproof-arbiter

It can read the job and critic result and write the final receipt.

It cannot rewrite what the critic observed.

When additional authority is genuinely required, the Arbiter can terminate with an explicit authority request rather than pretending verification succeeded.

Each job uses narrow, short-lived capabilities tied to one stable purpose:

agentproof-verify-before-commit

This means instructions hidden inside an untrusted candidate cannot simply grant themselves more permissions.


What AgentProof catches

I designed adversarial cases around failures that become dangerous once agents can act:

  • constraint violations
  • partial completion disguised as success
  • unsupported current-world claims
  • missing evidence
  • conflicting evidence
  • duplicate side effects
  • wrong target/entity for an action
  • authority gaps
  • prompt injection inside candidate data
  • ambiguous monetary values
  • incompatible currencies or physical units

For example:

“Product X is the cheapest option available today.”

without current supporting evidence cannot become SATISFIED.

AgentProof returns NEEDS_EVIDENCE.

Likewise, evidence that flight BB456 was already booked does not automatically block booking AA123. The action, target, and scope must actually match before AgentProof declares a duplicate side effect.


Reliability is part of the product

AgentProof is a service for other autonomous agents, so correctness under failure matters as much as the happy path.

The implementation includes:

  • idempotent requests
  • caller-isolated request identities
  • bounded concurrency and queueing
  • per-caller in-flight limits
  • fail-fast overload handling
  • whole-request timeouts
  • bounded model output
  • rate limits and payload limits
  • conservative retry behavior
  • bounded audit delivery
  • fail-closed SharedNet caller identity
  • auditable verification turns

One subtle example: if the caller times out but an external provider ignores cancellation, AgentProof keeps the idempotency key occupied until that underlying work actually terminates. A retry cannot quietly start a second copy of the same verification.


The hardest challenge

The hardest part was not detecting obvious failures.

It was making sure the verifier itself could not become overconfident.

During adversarial testing I repeatedly found cases where apparently sensible verification logic could still:

  • approve one part of a compound goal
  • compare unrelated dates
  • confuse two different actions or entities
  • compare numbers expressed in incompatible units
  • trust unsupported critic conclusions
  • return green while unresolved evidence conflicts remained
  • duplicate work after a timeout

Instead of patching only those examples, I kept moving the guarantees into structural invariants.

That became the most important lesson from building AgentProof:

Reliable agents need more than better reasoning. They need explicit contracts, bounded authority, provenance, independent verification, and honest uncertainty.


What I learned

SharedOS changed how I thought about agent reliability.

Permissions, purpose, escalation, and auditability are not secondary infrastructure concerns once agents can take consequential actions — they become part of the reasoning system itself.

AgentProof is my attempt to make that boundary reusable:

one small verification service that other agents can call before an uncertain output becomes a real-world action.

Verify before you trust.

Verify before you commit.

Built With

  • agentic-ai
  • agents
  • ai
  • api
  • devtools
  • distributed-systems
  • javascript
  • llm
  • node.js
  • openai-api
  • security
  • sharednet
  • sharedos
Share this project:

Updates

Submission history