Inspiration
Every freelancer and small-business owner knows this monthly ritual: a shoebox (or a phone gallery) full of receipts, a bank statement PDF, and an hour or two of squinting at both trying to figure out what matches what, what's a business expense, and what got missed. It's not hard work — it's just tedious, error-prone, and nobody wants to spend their evening doing it.
We wanted to build something that doesn't just automate this — it actually does the judgment work a bookkeeper would do, and only bothers you when something genuinely needs a human decision.
What it does
You give the agent two things: a folder of receipt images and a bank statement CSV. From there, it works entirely on its own:
- Reads each receipt and extracts the vendor, amount, and date
- Matches each receipt against the bank statement to confirm the payment actually happened
- Categorizes each expense for tax purposes
- Flags problems — a receipt with no matching bank transaction, a bank charge with no receipt, amounts that don't line up
- Produces a clean report, plus a short list of genuine questions for the user — e.g. "no receipt for this ₹499 charge — was this business or personal?"
The goal isn't to silently guess. It's to replace the manual matching work a person would do by hand, and hand back exactly the uncertain bits for a human to decide on.
How we built it
The agent is built on the Strands Agents SDK, with one orchestrator agent and five tools:
extract_receipt— reads a receipt image and pulls structured data (vision model)parse_bank_statement— parses the bank CSV into transactionsmatch_transaction— matches receipts to transactions using amount tolerance, date proximity, and fuzzy vendor-name matchingcategorize_expense— makes the tax-category judgment call, flagging genuine ambiguity honestly rather than guessing confidentlygenerate_report— compiles everything into a final reconciliation report
We deliberately kept matching and parsing as deterministic Python (fast, reliable, testable) and reserved model calls for the two steps that genuinely need judgment: reading receipt images and categorizing ambiguous expenses.
Challenges we ran into
The biggest challenge wasn't the agent logic — it was infrastructure. We originally built this against Amazon Bedrock, as the hackathon recommends. Partway through, the AWS account hit an unresolved model-access block (ValidationException: Operation not allowed) that persisted even after account verification succeeded, correct IAM/Marketplace permissions were confirmed, and a support case was filed and followed up on with AWS. We reproduced the issue with a minimal, isolated boto3 script to rule out our own code — the block was confirmed to be account-level, on AWS's side.
Rather than let that stall the project, we pivoted to running on Groq (OpenAI-compatible endpoint) instead. Because Strands is provider-agnostic by design, this was a contained change to the model configuration rather than a rewrite of the agent itself.
Along the way we also hit — and fixed — a string of smaller real-world problems: a reasoning model's multi-turn tool-calling conflicting with the OpenAI-compatible API (solved by passing images inline instead of via a tool call), Groq's daily token quota running out mid-build (solved by resizing images and splitting model usage across tools), and a receipt date-format bug where DD/MM/YY dates were being misread as MM/DD (solved with an explicit prompt instruction, after the agent itself flagged the resulting mismatch rather than silently mismatching it).
Accomplishments that we're proud of
- A genuinely autonomous pipeline — the agent decides its own tool-calling sequence rather than following a hardcoded script
- Watching it catch a real date-mismatch on an actual receipt and flag it as "likely the same transaction, worth confirming" instead of silently guessing wrong or dropping it — this is the core value proposition of the project, working exactly as intended
- Diagnosing and working around a genuine AWS account-level blocker without losing the build timeline, including filing and following up on a formal AWS support case
- A clean pivot from Bedrock to Groq that required changing model configuration only, thanks to Strands' provider-agnostic design
What we learned
- Agentic pipelines fail informatively, if you let them. The best moment in testing wasn't a clean run — it was watching the agent notice ambiguity and flag it rather than silently guess.
- Model-provider abstraction is worth the setup cost. Building on Strands meant the AWS→Groq pivot took an afternoon, not a rewrite.
- Deterministic code where you can, model calls where you must. Keeping matching/parsing as plain Python made debugging dramatically easier than if everything ran through an LLM.
What's next for ReceiptSync — Autonomous Reconciliation Agent
- Swap back to Amazon Bedrock / deploy via AgentCore if the account-access issue resolves
- Expand
categorize_expensewith a learnable per-user category mapping over time - Support PDF receipts natively (currently requires image conversion)
- Add a lightweight web UI so non-technical users don't need to run it from the terminal


Log in or sign up for Devpost to join the conversation.