ProofOps — Verified Autonomous Operations Agent

ATTEMPTED ≠ COMPLETED.

AI agents are becoming capable of taking real actions: sending emails, updating records, creating tickets, making bookings, and triggering business workflows.

But there is a reliability problem hidden behind that autonomy:

An API call being attempted does not prove that the intended real-world effect happened.

Suppose an agent sends an email through Gmail. Gmail successfully commits the message, but the network response is lost before the agent receives it.

What should the agent do?

If it assumes failure and retries, it may send the same message twice.

If it assumes success, it may claim something happened without evidence.

That failure mode inspired ProofOps.

ProofOps is a verified autonomous operations agent designed around one rule:

AI can propose an action. Deterministic systems must authorize, execute, verify, and prove it.


The Problem

Most agent workflows implicitly treat execution as something close to:

LLM decision → API call → success

Real systems are not that simple.

External actions can encounter:

  • lost responses after successful commits;
  • timeouts with unknown outcomes;
  • duplicate retry attempts;
  • stale or incomplete business data;
  • malformed AI output;
  • authorization ambiguity;
  • provider failures;
  • partial execution; and
  • situations where the system simply does not have enough evidence to know what happened.

For consequential operations, guessing is dangerous.

A system should not say "completed" merely because an API call returned—or because the AI believes it probably succeeded.


Our Solution

ProofOps treats verification as part of execution itself.

The workflow is:

Goal

↓

Trusted business evidence

↓

Gemini structured proposal

↓

Deterministic validation

↓

Human approval

↓

Durable execution state

↓

Provider action

↓

Reconciliation

↓

COMPLETED_VERIFIED

The key separation is intentional:

AI proposes. Deterministic systems authorize, execute, verify, and prove.

Gemini is used for structured reasoning and proposal generation, but the model does not decide whether a consequential action is authorized or whether an external effect should be considered complete.

Those decisions belong to deterministic control logic.


How ProofOps Works

1. Receive a Goal

The user provides an operational goal.

Our demonstration focuses on a synthetic renewal-outreach workflow where business evidence may lead to an email action.

2. Resolve Trusted Evidence

ProofOps gathers trusted business evidence from the connected system rather than allowing the model to invent operational facts.

For the demo, Google Sheets acts as a synthetic business-data source.

3. Generate a Structured AI Proposal

Gemini receives the trusted evidence and produces a typed, structured proposal.

The AI proposes what should happen.

It does not receive unrestricted authority to directly execute the consequential action.

4. Apply Deterministic Validation

Before execution, ProofOps validates the proposed operation against deterministic rules and expected schemas.

Malformed, unsafe, unsupported, or ambiguous proposals are rejected.

When the system cannot safely continue, it fails closed to:

BLOCKED_NEEDS_HUMAN

5. Require Human Approval

Consequential actions require explicit human approval.

The approved operation is tied to the exact action being reviewed rather than treating approval as a vague permission for whatever an agent decides to do later.

6. Create Durable Execution State

Before the external effect is attempted, ProofOps records durable execution state.

This provides a persistent source of truth even if the process crashes, a response disappears, or the provider becomes temporarily unavailable.

7. Perform the Provider Action

After authorization, the system performs the intended external action—for example, sending an email through Gmail.

8. Treat Uncertainty as a Real State

If the provider outcome cannot be proven, ProofOps does not guess.

The operation enters:

UNKNOWN

That distinction is critical.

UNKNOWN does not mean failure.

It means:

ProofOps does not currently have enough evidence to truthfully claim either success or failure.

9. Reconcile Before Retrying

Before another potentially duplicate side effect can occur, ProofOps examines provider state and searches for evidence of the original action.

If the external effect already happened, ProofOps binds the discovered provider evidence back to the logical operation instead of performing the effect again.

10. Verify Completion

Only after external evidence proves the intended result does the operation reach:

COMPLETED_VERIFIED

The system therefore distinguishes between:

  • wanting an action;
  • proposing an action;
  • approving an action;
  • attempting an action; and
  • proving that the intended real-world outcome exists.

Live Failure-Recovery Proof

The most important demonstration was not the normal success path.

We intentionally tested the dangerous case.

During fresh Gmail qualification:

  1. ProofOps initiated one logical email action.
  2. Gmail committed the physical message.
  3. The provider response was intentionally lost.
  4. ProofOps observed the outcome as UNKNOWN.
  5. It did not blindly resend the message.
  6. The reconciliation process searched provider state.
  7. It discovered the already-created message.
  8. The provider message ID was bound back to the logical effect.
  9. Verified evidence and a signed receipt were recorded.
  10. Exactly one physical Gmail message existed for the logical action.

The qualification demonstrated:

UNKNOWN_OBSERVED = TRUE
RECONCILED_WITHOUT_RESEND = TRUE
PHYSICAL_MESSAGES_FOR_LOGICAL_EFFECT = 1
PROVIDER_MESSAGE_ID_BOUND = TRUE
SIGNED_RECEIPT_PRESENT = TRUE

This is the behavior ProofOps was built to guarantee:

When reality is uncertain, investigate before acting again.


Safety Model

ProofOps follows several core safety rules:

Fail Closed

Unsafe, malformed, or ambiguous operations do not silently continue.

They transition to:

BLOCKED_NEEDS_HUMAN

Never Convert Uncertainty Into Success

Unknown provider outcomes remain:

UNKNOWN

until evidence resolves them.

Reconcile Before Another Effect

An uncertain external action must be investigated before the system considers performing the effect again.

Evidence Before Completion

COMPLETED_VERIFIED is reserved for operations whose intended outcome has been supported by external evidence.

Human Control for Consequential Actions

AI reasoning does not bypass deterministic authorization or human approval.

Auditable State

Important decisions, transitions, evidence, and execution outcomes are retained so the lifecycle of an operation can be inspected later.


How We Built It

ProofOps is built as a layered reliability architecture rather than a single autonomous LLM loop.

The backend uses Python 3.12+, FastAPI, Pydantic, Uvicorn, SQLAlchemy, and Alembic.

PostgreSQL provides durable production state, while an offline/demo configuration can use SQLite.

Gemini generates structured proposals from trusted evidence using Google's Gen AI tooling.

For real external integrations, ProofOps connects to:

  • Google Sheets API for synthetic business evidence;
  • Gmail API for external email effects, provider evidence, and reconciliation.

The user interface is implemented with HTML, CSS, and JavaScript.

The application is containerized with Docker and deployed on Railway.

Testing includes pytest, pytest-asyncio, Hypothesis, and HTTPX, covering both normal behavior and adversarial/error cases.

Database schema evolution is managed through Alembic migrations.


Architecture

                 ┌────────────────────┐
                 │     User Goal      │
                 └─────────┬──────────┘
                           ↓
                 ┌────────────────────┐
                 │ Trusted Evidence   │
                 └─────────┬──────────┘
                           ↓
                 ┌────────────────────┐
                 │ Gemini Proposal    │
                 │  Structured Only   │
                 └─────────┬──────────┘
                           ↓
                 ┌────────────────────┐
                 │ Deterministic      │
                 │ Validation         │
                 └─────────┬──────────┘
                           ↓
                 ┌────────────────────┐
                 │ Human Approval     │
                 └─────────┬──────────┘
                           ↓
                 ┌────────────────────┐
                 │ Durable Effect     │
                 │ State              │
                 └─────────┬──────────┘
                           ↓
                 ┌────────────────────┐
                 │ Provider Action    │
                 └─────────┬──────────┘
                           ↓
              ┌────────────┴────────────┐
              ↓                         ↓
       Known Outcome               UNKNOWN
              ↓                         ↓
       Verify Evidence            Reconcile
              │                         │
              └────────────┬────────────┘
                           ↓
                COMPLETED_VERIFIED

Engineering Validation

ProofOps RC11 reached:

402 / 402 deterministic tests passing

Additional qualification included:

  • high-risk repeat runs: 3 / 3 PASS;
  • order-dependence checks: PASS;
  • deployed authentication boundary: PASS;
  • malformed and oversized input boundaries: PASS;
  • receipt tamper rejection: PASS;
  • duplicate JSON rejection: PASS;
  • fresh PostgreSQL live qualification: PASS;
  • fresh Google Sheets live qualification: PASS;
  • fresh Gmail live qualification: PASS.

The final formal Gemini live qualification was blocked by external provider/quota availability during qualification.

We deliberately did not turn that into a PASS.

That behavior reflects the philosophy of the project itself:

If something has not been proven, ProofOps does not claim that it has.


Challenges We Faced

1. The Hardest Failure Is Not Always "Failure"

Traditional error handling usually asks whether an operation succeeded or failed.

Distributed systems introduce a third possibility:

we do not know.

Designing UNKNOWN as a first-class execution state changed how the entire workflow had to behave.

2. Preventing Duplicate External Effects

Retrying HTTP requests is easy.

Retrying real-world actions safely is not.

A send operation cannot simply be repeated because the previous request timed out. We needed durable logical-effect identity and provider reconciliation before another side effect could be considered.

3. Separating AI Reasoning From Authority

LLMs are useful for interpreting context and proposing actions, but operational authorization needs stronger guarantees.

We therefore separated probabilistic reasoning from deterministic execution control.

4. Proving Completion

A successful function return is not sufficient evidence.

We had to define what external evidence is strong enough to move an operation into COMPLETED_VERIFIED.

5. Testing Failure Paths

Happy-path demos are easy to make impressive.

The difficult work was deliberately testing malformed data, duplicate input, uncertain provider responses, authentication boundaries, tampered receipts, repeat executions, and recovery behavior.


What We Learned

The biggest lesson from building ProofOps was that agent reliability is not only a model-quality problem.

A more intelligent model does not solve uncertainty after an external side effect.

We learned that reliable autonomous systems need:

  • durable execution state;
  • explicit uncertainty;
  • deterministic authorization;
  • exact human approval boundaries;
  • idempotent logical operations;
  • provider reconciliation;
  • external verification; and
  • auditable evidence.

Most importantly:

Verification should not be an optional step after execution. Verification is part of execution.


Why ProofOps Matters

As autonomous agents move from answering questions to changing the real world, the cost of an incorrect assumption increases.

The same uncertainty pattern demonstrated with Gmail can appear in:

  • payments;
  • bookings;
  • CRM updates;
  • ticket creation;
  • cloud operations;
  • database mutations; and
  • other consequential workflows.

ProofOps does not attempt to solve every possible provider integration today.

Instead, it demonstrates a reusable reliability principle:

Do not increase an agent's authority faster than its ability to prove what actually happened.

ProofOps treats uncertainty as state that must be resolved—not hidden.


Release

ProofOps RC11 — 1.0.0rc11

All demo and qualification workflows use synthetic/test identities.

ProofOps — the truth layer for autonomous agents that act in the real world.

Built With

Share this project:

Updates

Submission history