Inspiration

Every freelancer and small-business owner knows this monthly ritual: a shoebox (or a phone gallery) full of receipts, a bank statement PDF, and an hour or two of squinting at both trying to figure out what matches what, what's a business expense, and what got missed. It's not hard work — it's just tedious, error-prone, and nobody wants to spend their evening doing it.

We wanted to build something that doesn't just automate this — it actually does the judgment work a bookkeeper would do, and only bothers you when something genuinely needs a human decision.

What it does

You give the agent two things: a folder of receipt images and a bank statement CSV. From there, it works entirely on its own:

  1. Reads each receipt and extracts the vendor, amount, and date
  2. Matches each receipt against the bank statement to confirm the payment actually happened
  3. Categorizes each expense for tax purposes
  4. Flags problems — a receipt with no matching bank transaction, a bank charge with no receipt, amounts that don't line up
  5. Produces a clean report, plus a short list of genuine questions for the user — e.g. "no receipt for this ₹499 charge — was this business or personal?"

The goal isn't to silently guess. It's to replace the manual matching work a person would do by hand, and hand back exactly the uncertain bits for a human to decide on.

How we built it

The agent is built on the Strands Agents SDK, with one orchestrator agent and five tools:

  • extract_receipt — reads a receipt image and pulls structured data (vision model)
  • parse_bank_statement — parses the bank CSV into transactions
  • match_transaction — matches receipts to transactions using amount tolerance, date proximity, and fuzzy vendor-name matching
  • categorize_expense — makes the tax-category judgment call, flagging genuine ambiguity honestly rather than guessing confidently
  • generate_report — compiles everything into a final reconciliation report

We deliberately kept matching and parsing as deterministic Python (fast, reliable, testable) and reserved model calls for the two steps that genuinely need judgment: reading receipt images and categorizing ambiguous expenses.

Challenges we ran into

The biggest challenge wasn't the agent logic — it was infrastructure. We originally built this against Amazon Bedrock, as the hackathon recommends. Partway through, the AWS account hit an unresolved model-access block (ValidationException: Operation not allowed) that persisted even after account verification succeeded, correct IAM/Marketplace permissions were confirmed, and a support case was filed and followed up on with AWS. We reproduced the issue with a minimal, isolated boto3 script to rule out our own code — the block was confirmed to be account-level, on AWS's side.

Rather than let that stall the project, we pivoted to running on Groq (OpenAI-compatible endpoint) instead. Because Strands is provider-agnostic by design, this was a contained change to the model configuration rather than a rewrite of the agent itself.

Along the way we also hit — and fixed — a string of smaller real-world problems: a reasoning model's multi-turn tool-calling conflicting with the OpenAI-compatible API (solved by passing images inline instead of via a tool call), Groq's daily token quota running out mid-build (solved by resizing images and splitting model usage across tools), and a receipt date-format bug where DD/MM/YY dates were being misread as MM/DD (solved with an explicit prompt instruction, after the agent itself flagged the resulting mismatch rather than silently mismatching it).

Accomplishments that we're proud of

  • A genuinely autonomous pipeline — the agent decides its own tool-calling sequence rather than following a hardcoded script
  • Watching it catch a real date-mismatch on an actual receipt and flag it as "likely the same transaction, worth confirming" instead of silently guessing wrong or dropping it — this is the core value proposition of the project, working exactly as intended
  • Diagnosing and working around a genuine AWS account-level blocker without losing the build timeline, including filing and following up on a formal AWS support case
  • A clean pivot from Bedrock to Groq that required changing model configuration only, thanks to Strands' provider-agnostic design

What we learned

  • Agentic pipelines fail informatively, if you let them. The best moment in testing wasn't a clean run — it was watching the agent notice ambiguity and flag it rather than silently guess.
  • Model-provider abstraction is worth the setup cost. Building on Strands meant the AWS→Groq pivot took an afternoon, not a rewrite.
  • Deterministic code where you can, model calls where you must. Keeping matching/parsing as plain Python made debugging dramatically easier than if everything ran through an LLM.

What's next for ReceiptSync — Autonomous Reconciliation Agent

  • Swap back to Amazon Bedrock / deploy via AgentCore if the account-access issue resolves
  • Expand categorize_expense with a learnable per-user category mapping over time
  • Support PDF receipts natively (currently requires image conversion)
  • Add a lightweight web UI so non-technical users don't need to run it from the terminal

Built With

  • aws-bedrock
  • aws-iam
  • boto3
  • groq
  • openai
  • pandas
  • pillow
  • python
  • qwen
  • rapidfuzz
  • strands-agents-sdk
Share this project:

Updates

Submission history