Inspiration

A field-service job isn't finished when the work is finished. It's finished when someone in an office goes through a pile of photos, receipts and signatures, decides whether they actually prove the work order was met, chases the technician for the photo they forgot, and only then sends the invoice. That review is slow, boring and error-prone, and it's why invoices go out days late.

It's also the kind of work agents are supposed to be good at - except getting it wrong costs money in both directions. Close a job that isn't really done and you invoice for work nobody can prove. Refuse one that is done and you hold up payment and annoy a technician who did everything right. That tension is what made it worth building: an agent trustworthy enough to actually let go of.

What it does

A technician uploads whatever evidence the job produced - photos, a receipt, a signature, a voice note - and leaves. FieldProof then:

  • reads every artifact into structured observations;
  • checks each resulting claim against the work order, including whether the evidence is even capable of supporting it (a receipt proves purchase, never installation);
  • asks the technician once, in a single message, for anything missing or unreadable;
  • escalates to a supervisor only when a decision genuinely needs a human, as one question with the number in it: "Approve one additional filter at $54 outside the original scope?";
  • generates the closeout report, invoices, notifies the customer, closes the job and seals a SHA-256 receipt over the whole evidence chain.

On our reference scenario that's 18 workflow steps, 16 of them with no human involved: one text to the technician, one decision for the supervisor, zero manual document review.

How we built it

Strands Agents SDK on Amazon Bedrock, with a deliberate line between what a model decides and what code decides. Most agent demos let the model decide. When the decision moves money, that's a bug - so our models read evidence and write messages, and nothing else.

  • Five agents: Context, Evidence, Reconciliation, Policy, Action. The Evidence Agent is the multimodal one - each artifact type has a primary reader and an alternate, and anything unreadable counts as nothing rather than as present.
  • A Strands multi-agent Graph for orchestration: seven nodes, conditional edges off the policy verdict, node and execution timeouts. The nodes are custom MultiAgentBase nodes, because every decision - claim construction, evidence compatibility, reconciliation, policy verdicts, authorization, state transitions - is deterministic code. It's the hybrid the Strands graph docs describe: model nodes for judgement, deterministic nodes for control.
  • A Strands @tool surface so a model can drive a closeout itself through the SDK's tool loop - safe because authorization isn't in the prompt.
  • Deployed two ways: Bedrock AgentCore Runtime, or a SAM stack with DynamoDB, S3, EventBridge and two Lambdas. Same workflow code on both.

Challenges we ran into

Conflict identity. Reconciliation recomputes from scratch on every pass, so a conflict has to be recognisable as the same problem across passes - otherwise a supervisor mid-decision gets a duplicate question the moment new evidence lands. Conflict ids became a fingerprint of what the conflict is about, values included, so an approval for 3-against-2 can't silently carry over to 4-against-2.

Keeping the model out of the decisions without making it useless. The tempting shortcut is to let the agent judge whether a requirement is met. That's exactly the judgement that has to be reproducible, so it became a compatibility matrix and a rules table - and the model got the work it's genuinely better at.

Failing closed without failing shut. An unreadable artifact must not verify anything, but it also must not dead-end the job. It becomes a confidence-0.0 observation, which reconciliation turns into a specific request for a clearer copy.

Accomplishments that we're proud of

You can't prompt your way past the policy layer. Hand the closeout tools to a model and ask it to close an unfinished job, and it gets back ok=false with the invariant that stopped it - INV-001 - and the job doesn't move. Handing an agent real authority is only interesting if something can tell it no.

Ten invariants enforced in code and pinned by 113 tests, including the whole 17-scenario evaluation dataset run through both orchestrators.

Pauses are real stops. Waiting on a human ends the run; the next event starts a new one. Nothing sleeps, so resumption survives a restart - and at-least-once delivery is safe, because every side effect carries an idempotency key, a refused action doesn't burn its key, and a failed one releases it.

What we learned

That an agent framework earns its place at the edges of a system rather than the middle. Strands made the model work - multimodal reads, structured output, tool calling, the graph runtime, AgentCore deployment - small enough that the interesting engineering could go into the part that decides. The measure of a good agent system turned out to be how little of it the model is trusted with.

What's next for FieldProof

Real invoicing (QuickBooks or Stripe behind the existing submit() port), an identity provider in place of the API key, per-customer policy packs, and a mobile capture flow that nudges the technician before they leave site rather than after.

Built With

Share this project:

Updates

Submission history