WorkProof Sentinel — Don't Trust “Done.” Verify It.

Inspiration

In professional engineering and vendor delivery, “Done” is often a status—not proof.

A Jira ticket can say DONE. A GitHub pull request can say MERGED. A vendor can submit a $120,000 migration invoice.

But none of those statements prove that the underlying work actually exists, is healthy, and satisfies the original requirements.

We asked:

What if an AI agent didn't trust completion claims, but actively tried to disprove them?

That question became WorkProof Sentinel.

WorkProof is an autonomous evidence-verification agent that turns professional completion claims into falsifiable requirements, gathers evidence across business and infrastructure systems, searches for contradictions and missing evidence, actively attempts to invalidate its own conclusion, and escalates only when human judgment is genuinely required.


What WorkProof Does

WorkProof follows a simple principle:

Claim → Requirements → Evidence → Counter-Evidence → Verdict

Given a professional task, SOW, ticket, or delivery claim, WorkProof:

  1. Decomposes the claim into atomic verification requirements.
  2. Plans evidence queries across multiple systems.
  3. Queries infrastructure and business data using Steampipe-style relational SQL.
  4. Checks temporal consistency between migrations, tests, backups, approvals, and sign-offs.
  5. Detects contradictions between systems such as Jira, GitHub, AWS RDS, Backup, and CloudWatch.
  6. Actively falsifies the claim by searching for evidence that would prove the completion statement wrong.
  7. Refuses certification when required evidence is missing or contradictory.
  8. Escalates to a human only when an actionable evidence decision is required.
  9. Resumes from checkpoint after new evidence is supplied.
  10. Produces a tamper-evident HMAC-SHA256 receipt containing the final verification result and evidence provenance.

The important behavior is not that WorkProof says “yes.”

It is that WorkProof is willing to say:

“No. This is not proven yet.”


The AWS Agent Architecture

WorkProof is built around the AWS Strands Agents SDK, with Amazon Bedrock used for live reasoning and AgentCore-oriented state/checkpoint handling.

The agent exposes a verification decision loop:

Thought → Action → Observation → Decision

Instead of blindly executing a fixed checklist, the verification layer records which evidence is being investigated, what was observed, and why the next verification action is required.

The original verification intelligence includes:

  • SOW Claim Decomposer
  • Evidence Query Planner
  • Temporal Consistency Engine
  • Cross-Source Contradiction Detector
  • Adversarial Falsification Engine
  • Evidence Sufficiency Engine
  • Checkpoint / Resume State
  • Tamper-Evident Receipt Generator

This gives the agent a fundamentally different objective from a normal assistant:

Normal assistant: “Can I explain why this looks complete?”

WorkProof: “What evidence would prove this claim wrong?”


The 90/10 WIN SMART Architecture

We deliberately did not rebuild API connectors, authentication layers, database drivers, and SaaS integrations from scratch.

90% — Mature Open-Source Infrastructure

WorkProof uses Steampipe as the multi-system SQL layer when live execution is available.

This allows the verification layer to reason over relational representations of systems such as:

  • AWS
  • GitHub
  • Jira
  • CloudWatch
  • AWS Backup

Instead of writing separate API wrappers for every system, verification logic can operate through SQL queries and joins.

For example:

SELECT i.db_instance_identifier,
       i.status,
       b.status AS backup_status
FROM aws_rds_db_instance i
JOIN aws_backup_recovery_point b
  ON b.resource_name = i.db_instance_identifier;

Dual-Engine Evaluation

To make the project immediately runnable without AWS credentials, SaaS tokens, or local daemons, WorkProof has two execution modes:

LIVE_STEAMPIPE_CLI

Uses an installed Steampipe CLI and its configured plugins.

BUNDLED_REPLAY_MIRROR

Uses a deterministic relational replay environment containing the same verification-oriented table structures and realistic evidence records.

The UI and API explicitly report which execution mode is active.

This makes the project both production-connectable and zero-friction to evaluate.


The Killer Demo: The $120,000 RDS Migration

A vendor claims that four production database clusters have been completely migrated:

  • db-users
  • db-orders
  • db-billing
  • db-analytics

The supporting systems appear convincing.

Jira: DONE GitHub: PR #412 MERGED Invoice: $120,000

A normal workflow might approve the delivery.

WorkProof investigates.

Step 1 — Decompose

The agent converts the delivery claim into atomic requirements.

REQ-01: Four production clusters must exist
REQ-02: Migration must have completed
REQ-03: Required validation evidence must exist
REQ-04: No blocking contradictions may remain

Step 2 — Gather Evidence

WorkProof queries the available evidence sources.

Step 3 — Cross-Check

Jira and GitHub indicate completion.

But infrastructure evidence tells a different story.

Step 4 — Attack Its Own Conclusion

The adversarial verifier asks:

“What evidence would make the completion claim false?”

The search finds that the expected fourth cluster, db-analytics, does not have the required validation evidence.

Step 5 — Refuse Certification

Instead of producing a reassuring summary:

VERDICT: NOT PROVEN

Reason:
Required production evidence is missing.

Human action:
Provide the missing validation artifact.

The agent checkpoints the audit state.

Step 6 — Human Evidence Injection

The engineer supplies the missing evidence.

WorkProof restores the checkpoint and continues verification rather than restarting the entire audit.

Step 7 — Final Certification

The evidence is re-evaluated.

The falsification checks pass.

VERDICT: PROVEN
Requirements: 4/4
Contradictions: 0
Falsification: PASSED

WorkProof generates a tamper-evident HMAC-SHA256 receipt containing the resulting proof package.


Why This Matters

WorkProof targets a recurring professional problem:

Organizations have enormous amounts of system data, but completion decisions are still often based on status fields, screenshots, messages, and trust.

This matters for:

  • Vendor delivery verification
  • Cloud migration sign-off
  • Infrastructure change certification
  • Compliance evidence collection
  • SOC2-style operational checks
  • Engineering acceptance testing
  • Professional service deliverables
  • Invoice and milestone verification

The goal is not to replace the responsible engineer or manager.

The goal is to make the human decision happen after the evidence has been assembled and challenged, rather than before.


Challenges We Solved

Cross-System Consistency

GitHub, Jira, AWS, and telemetry systems represent time and state differently. WorkProof normalizes timestamps and evaluates causal relationships instead of trusting isolated status fields.

Adversarial Verification

An LLM naturally tends to rationalize the evidence it receives.

We explicitly designed the verification layer to search for counter-evidence, including missing resources, inconsistent statuses, incomplete validation artifacts, and temporal anomalies.

Reliable Human Escalation

When evidence is insufficient, the agent does not simply fail.

It creates an actionable checkpoint, waits for the missing decision or artifact, and resumes from the preserved verification state.

Zero-Friction Evaluation

Live infrastructure integrations are useful in production, but they make hackathon evaluation difficult.

The dual-engine design lets judges run a deterministic replay locally while still exposing the live Steampipe execution path.


Technical Implementation

Core stack:

  • AWS Strands Agents SDK
  • Amazon Bedrock
  • AgentCore-oriented checkpoint/state handling
  • Steampipe OSS
  • Python 3.12
  • FastAPI
  • Uvicorn
  • SQL / relational evidence model
  • Pytest / Pytest-Cov
  • HMAC-SHA256

The project includes automated verification for the agent workflow, evidence queries, contradiction detection, temporal checks, adversarial falsification, REST APIs, and cryptographic receipt verification.

Current test status:

30 tests passing — 93% code coverage


What Makes WorkProof Different?

Most enterprise agents are optimized around action:

“Do the task.”

WorkProof is optimized around verification:

“Prove the task was actually done.”

And unlike a passive compliance checklist, it does not only collect supporting evidence.

It deliberately searches for evidence that could make its own conclusion wrong.

That creates a useful safety property:

A completion claim must survive an attempt to falsify it before WorkProof will certify it.


What's Next

The architecture can expand beyond infrastructure migration verification into:

  • Multi-cloud verification
  • Automated remediation playbooks
  • Compliance policy ingestion
  • Vendor milestone verification
  • Contract/SOW acceptance workflows
  • Additional Steampipe data sources
  • Automated evidence collection and remediation

The long-term vision is simple:

Make “Done” an evidence-backed state, not a checkbox.


Built With

  • amazon-bedrock
  • aws-strands-sdk
  • fastapi
  • hmac-sha256
  • playwright
  • postgresql
  • pytest
  • python
  • steampipe
  • tailwindcss
  • uvicorn
Share this project:

Updates

Submission history