Inspiration

Most agent demos declare success when an API returns 200 OK. Real professional operations are messier: records change while approvals are pending, networks fail after writes, and a retry can duplicate an expensive action.

I built Proofline because support and operations teams need AI that can clear repetitive cross-system work without taking judgment away from people—or claiming success it cannot prove.

What it does

Proofline Strands coordinates a customer refund across Salesforce, Slack, and Stripe Test Mode.

It turns each request into an immutable Outcome Contract containing the case, payment, amount, currency, reason, and current Salesforce revision. That contract is hashed into a canonical SHA-256 action digest.

A human approves that exact digest in Slack. Before executing anything, Proofline independently verifies the approval and confirms that the Salesforce record has not changed. Only then can its deterministic safety kernel perform one idempotent Stripe Test Mode refund.

Proofline does not consider the workflow complete merely because a tool call returned successfully. It reads Stripe and Salesforce back independently and produces PROVEN only when Slack, Stripe, and Salesforce all agree.

The reliability test

The demo deliberately recreates a dangerous real-world failure: Stripe commits the refund, but its response disappears.

A typical agent might retry the payment operation or confidently report success. Proofline does neither. It moves to UNKNOWN, withholds the completion claim, and searches Stripe's authoritative state using immutable run metadata.

It finds the original refund, proves that exactly one mutation occurred, reconciles Salesforce, and independently reads the updated record back. Only then does the verdict become PROVEN.

How I built it

  • AWS Strands Agents SDK for TypeScript provides structured intent extraction, reasoning, and guarded tool selection.
  • Five narrow Strands tools expose verification, test-record creation, approval requests, guarded execution, and evidence inspection.
  • A deterministic state machine controls authorization and legal transitions outside the model.
  • Slack provides the human judgment point.
  • Stripe Test Mode provides the consequential financial write.
  • Salesforce Developer Edition provides the business record and reconciliation receipt.
  • Amazon Bedrock AgentCore Runtime hosts the Node.js 22 agent using the same HTTP application that runs locally.
  • AWS CloudWatch provides runtime logs and operational visibility.
  • Eight executable safety tests cover digest tampering, stale records, illegal transitions, missing evidence, tool boundaries, and ambiguous writes.

The model never receives a tool that can directly call Stripe. Every consequential execution must pass through the deterministic approval and evidence kernel.

Challenges I faced

The hardest problem was not connecting three APIs. It was defining what “done” means when each application may expose a different truth.

I had to separate proposal, authorization, execution, observation, reconciliation, and proof. Approval also needed to remain valid only for the exact action and Salesforce revision the human reviewed.

Deploying to AgentCore introduced additional production constraints, including Node.js 22 bundling, a read-only /var/task filesystem, writable /tmp evidence storage, and AgentCore's health and invocation contracts.

What I learned

Human-in-the-loop is strongest when the human approves an immutable action rather than a vague natural-language intention.

I also learned that idempotency prevents duplicate writes, but it does not prove the business outcome. Reliable agents need explicit observation, reconciliation, and an honest UNKNOWN state.

The model should propose and coordinate. Deterministic controls should govern authority. Independent evidence should determine completion.

Accomplishments I am proud of

  • A working workflow across three official vendor test/developer environments.
  • Human approval cryptographically bound to the precise action.
  • Recovery from a lost response without performing a second mutation.
  • Eight passing executable safety tests.
  • A success claim that is mechanically unreachable until all three systems agree.
  • Successful deployment to Amazon Bedrock AgentCore Runtime.

What's next

Proofline's safety kernel is operation-agnostic. Future policy packs could support credits, invoice adjustments, subscription changes, procurement approvals, and account restoration.

I would also add durable evidence storage, AgentCore Identity for managed third-party credentials, and reusable Outcome Contract evaluators for other Strands agents.

AI can propose. Proofline makes it prove.

Built With

Share this project:

Updates

Submission history