Inspiration

Accounts-Payable and procurement teams spend thousands of hours manually cross-checking invoices against purchase orders, contract pricing rules, and warehouse receipts. Traditional software stops at basic OCR field extraction, leaving human teams to manually investigate missing data, calculate price/quantity variances, and detect duplicate billing. We were inspired to build a true autonomous procurement agent using the Strands Agents SDK that performs multi-evidence correlation, reasoning, anomaly risk assessment, and decision routing end-to-end.

What it does

InvoiceGuard is an autonomous procurement agent that ingests vendor invoice PDFs, investigates procurement evidence across Purchase Orders (POs), Master Contracts, Goods Receipt Notes (GRNs), and historical anti-fraud data, calculates a transparent deterministic risk score (0–100), auto-approves low-risk routine invoices, and prepares actionable human escalation workflows (such as drafting vendor clarification notices or escalating to procurement managers) only when human decision-making is genuinely required.

How we built it

  • Agent Core: Built with the Strands Agents SDK (strands.Agent and @tool primitives) orchestrating tools for PDF extraction, PO matching, contract term lookup, GRN quantity verification, and historical duplicate detection.
  • Backend Service: FastAPI REST API, Python 3.10+, SQLAlchemy, and SQLite database.
  • Frontend Dashboard: Polished React + Vite interface with Google Fonts (Outfit / Inter), dark glassmorphism styling, visual risk score gauge, live "Agent Thinking" action timeline, correlated evidence cards, human decision controls, and an immutable audit log viewer.
  • Testing & Evaluation: Synthetic PDF generator using ReportLab and an automated evaluation harness (scripts/evaluate_agent.py) benchmarking decision accuracy across 7 hackathon test scenarios.

Challenges we ran into

  • Preventing Arithmetic Hallucinations: LLMs can occasionally miscalculate percentage variances or boundary conditions. We solved this by pairing the Strands Agent's reasoning and tool selection with deterministic Python code for exact variance math and risk formulas.
  • Multi-Evidence Data Correlation: Mapping unstructured text extracted from PDFs to relational database models across varying vendor formatting while ensuring historical duplicate detection ignored normal recurring monthly payments.

Accomplishments that we're proud of

  • 100% Decision Accuracy: Achieved a 100% pass rate on our evaluation harness (evaluate_agent.py) across all 7 hackathon test scenarios (clean invoices, PO price mismatches, quantity discrepancies, contract violations, duplicate invoices, and missing POs).
  • Sub-Second Processing Latency: Average processing speed of ~0.066 seconds per invoice while executing 7 tool calls per run.
  • Zero-Cost Local Execution: Designed a architecture that runs 100% locally out-of-the-box using free open-source components, while remaining fully ready for optional Amazon Bedrock and AgentCore cloud deployment.

What we learned

  • How to leverage Strands Agents SDK @tool decorators to separate semantic decision-making from deterministic business calculations.
  • How to design human-in-the-loop UX that presents structured evidence cards and actionable recommendations rather than raw chain-of-thought outputs.

What's next for InvoiceGuard

  • Live ERP Connectors: Building native integrations for SAP S/4HANA, NetSuite, and QuickBooks Online.
  • Amazon Bedrock AgentCore Scaling: Deploying the Strands agent to Amazon Bedrock AgentCore Runtime for enterprise multi-tenant serverless execution.
  • Automated AP Inbox Assistant: Monitoring corporate accounts-payable email aliases to automatically extract, verify, and resolve vendor invoice exceptions in real time.

Built With

Share this project:

Updates

Submission history