Inspiration

Small design and software studios constantly do extra work that never gets billed: one more page, a rush change, a third revision round the agreement didn't cover. The evidence is scattered across the agreement, client messages, approvals and delivery events, and chasing it feels awkward. The obvious AI answer, "let a model read the inbox and invoice the client", is exactly the wrong one: a model that can be talked into believing a client approved something can be talked into charging them.

What it does

Earned watches a project's agreement, client messages, approvals, deliveries and invoices. When approved, delivered work is missing from an invoice, it shows the owner one decision with every link in the evidence chain cited to its source with a verbatim quote. On approval it issues the invoice, sends it, and reconciles the payment. When the evidence is incomplete, disputed or suspicious, it says exactly what is missing and bills nothing.

  • Agreement intake: paste an agreement; a Strands agent extracts price, revision rounds and billing trigger, each with a verbatim quote. Only an owner-confirmed agreement can authorise billing.
  • Client portal: the client accepts a quote or delivery on an authenticated channel, creating real acceptance evidence.
  • Works in the background: a sweep re-checks open cases, sends at most two polite reminders (never in quiet hours, never after a dispute or payment) and emails the owner one digest only when a decision is waiting.
  • Guardrails you can see: every tool call the follow-up agent makes, allowed or denied, is logged with the policy reason.
  • Honest money: four separate states (possible, approved-unbilled, issued, paid), never summed, never the same rupee in two places.

How we built it

The core rule: a model reads the evidence; code decides whether money can be charged. The model can always point at a reason not to charge; it can never create a reason to charge on its own word. That is enforced three independent times:

  1. Verbatim grounding: every claim must quote its source exactly, or it is rejected. A paraphrase or invented quote never reaches a decision.
  2. Billing policy: charge-enabling facts count only when the evidence arrived through an authenticated channel from a person the owner configured with that authority. Amounts are parsed from text in code, in integer minor units; the model never does arithmetic.
  3. Cedar policy at the tool gateway: the follow-up agent's side-effect tools (send_reminder, create_payment_link) are authorised by real Cedar policies evaluated outside the agent's code. A reminder to the wrong recipient, a payment link for a tampered amount, or a reminder over the cap is denied deterministically. Locally via cedarpy; the same .cedar files target an Amazon Bedrock AgentCore Gateway.

The pipeline is a stateless Strands Agents SDK Graph (evidence_reader agent → deterministic ground node → deterministic decide node → brief_writer agent), with structured output and one bounded retry for rejected quotes. Snapshot in, record out, so the same graph runs in-process and on Amazon Bedrock AgentCore Runtime. FastAPI serves the API and a plain HTML/CSS/JS owner UI; SQLite holds tenant-scoped business state. Owner approval binds to a payload hash and evidence version (stale approvals return HTTP 409), and unique allocations prevent double issue.

Challenges we ran into

  • Making grounding strict (Unicode and whitespace normalisation only) without making honest extractions fail. The fix is always the prompt, never loosening the check.
  • Separating "did the model read this or make it up?" (grounding) from "does this person have authority?" (policy). Collapsing the two is how spoofed approvals get through.
  • Getting live Amazon Bedrock access in time. The demo and evaluation therefore run in replay mode (recorded model output) with simulated email and payments, and that is labelled everywhere it appears, in the UI and in the video.

Accomplishments that we're proud of

  • A spoofed "I accept, invoice me now" email is extracted, then shown as not trusted, and nothing is billed.
  • 17 decision-layer adversarial cases pass through the real graph with 0 unauthorised charges, and 119 automated tests cover policy, grounding, the Strands graph, the Cedar gateway and API flows.
  • The guardrail is something you can watch work, not just trust.

What we learned

Agents for humans earn trust by being boring where it matters. Let the model do what it is good at (reading messy evidence) and put every money decision behind deterministic, testable code and policy that no prompt can talk its way past.

What's next for Earned

Live Bedrock evaluation with every attempt reported, deployment to AgentCore Runtime with Cedar policies on an AgentCore Gateway, real email through Amazon SES, and Razorpay payment links for studios in India.

Built With

Share this project:

Updates

Submission history