Inspiration

Most invoice packets are routine, but someone still has to open them. I built PayablePilot to finish the clean work and stop at the one decision an agent should not make by itself: whether to pay a price exception.

It is for accounts payable reviewers, controllers, and small finance teams. Their bottleneck is not adding numbers. It is checking three documents, finding the exception, gathering context, and leaving a clean trail for the person who owns the decision.

What it does

The demo has two fictional invoice packets. PP-2087 clears because its purchase order, invoice, and receipt agree. PP-2086 stops because eight monitor arms were invoiced at $119 each while the purchase order lists $94. The difference is $25 per unit and $200 in total.

The Strands agent works through six tool calls. It lists the queue, inspects both packets, checks the exception supplier with live SerpApi data, clears the clean packet, and sends the exception to a person. CDW Canada is a real supplier reference. The purchase order and invoice are clearly marked demo records.

The live supplier check matched CDW Canada across three identity sources and found no adverse news tied to that company. That evidence is context, not a fraud score. A reviewer can request a credit and hold the invoice or choose to pay it anyway. The agent cannot make that choice.

How I built it

I used the Strands Agents SDK for the agent loop and tool selection. A separate TypeScript layer reads the packet fixture, checks the three documents, and calculates the price difference. The agent decides which tool to call next. Ordinary code owns the quantities, prices, and final dollar amount.

The repository also includes an AgentCore-compatible HTTP runtime and a deterministic test model that exercises the real Strands loop without cloud credentials. Fourteen tests cover the agent, financial checks, supplier evidence, HTTP invocation contract, and API response.

The public page replays the included fixture so judges can inspect both states without AWS credentials. It is not a live accounting connection. The SerpApi view is backed by a fresh run from August 30. One public evidence file records the six-tool sequence, both SerpApi search IDs, and the verified $200 calculation. A second proves the local /ping and /invocations contract. Neither file claims an AWS deployment or Bedrock run.

Challenges

The hard part was deciding where the agent had to stop. Letting a model calculate the amount would make the demo shorter, but it would also make the result harder to audit. I kept every money calculation in the domain layer and limited the agent to choosing bounded tools.

The supplier check needed the same restraint. A missing result is reported as unverified. Search evidence never becomes an invented risk label.

What is working

  • PP-2087 clears only after the three documents agree.
  • PP-2086 is held on a verified $200 price difference.
  • The public proof records all six tool calls.
  • The supplier check returned two SerpApi search IDs.
  • The live search matched CDW Canada across three identity sources.
  • A person must make the exception decision.
  • All 14 tests pass.

Why it matters

A finance team should not have to choose between slow manual review and an agent that improvises with money. PayablePilot automates the repetitive path, keeps arithmetic deterministic, and gives the reviewer one specific decision with the evidence attached.

What I learned

The useful part of this agent is not financial improvisation. It is choosing the next bounded action while ordinary code keeps the facts and arithmetic stable.

What is next

  • Connect the runtime to a real invoice queue
  • Persist tool calls and review decisions
  • Extract purchase orders and invoices instead of using fixtures
  • Add customer-specific approval policies

Built With

  • accounts-payable
  • nextjs
  • serpapi
  • strands-agents
  • typescript
Share this project:

Updates

Submission history