Inspiration

Akshaya Patra feeds 2.35 million kids every day. That number is hard to picture. But behind every meal is a pile of paper.

To get government subsidies, they follow FCRA rules. Every rupee spent in the field has to be labeled either "programmatic" (feeding kids) or "administrative" (office, salaries, general ops). And administrative spending can't cross 20% of foreign funding. That's a legal line you do not cross.

Here's the problem. Field coordinators buy petrol and vegetables with cash. They get handwritten receipts in Hindi or Tamil. Sometimes a thumbprint instead of a signature. And a lot of these expenses are mixed. A single ₹500 petrol receipt might cover three school visits (programmatic) AND one trip to a government office (administrative).

Right now, an accountant has to sit down and split every single one by hand. Get it wrong, and funding gets delayed. Which means kids don't get lunch tomorrow.

We built ComplyMate because that felt like something an AI agent could actually help with — not replace the accountant, but handle the boring parts so they can focus on the calls that need a human. What it does

ComplyMate reads messy field receipts and turns them into FCRA-compliant classifications.

For each receipt, it:

  • Pulls out the amount, vendor, date, and any handwritten notes
  • Checks the field coordinator's travel log for that day
  • Applies the FCRA rules to split the expense
  • Checks whether the 20% cap is still safe
  • Pauses and asks a human when something is unclear, missing, or risky
  • Writes the final decision into an audit ledger

Say a receipt comes in for ₹500 of petrol. The notes say "3 school visits + 1 govt office." The agent splits it into ₹375 programmatic and ₹125 administrative, checks that this keeps the total admin spend at 16.2% (safe), and asks the accountant to confirm with one click.

If the split would push admin spending over 20%, the agent stops. It says HARD STOP and refuses to record the expense. It doesn't guess. It doesn't quietly approve. It asks.

How we built it

We used the AWS Strands Agents SDK to build the agent, and Groq (running Llama 4 Scout) as the model. The whole thing runs in Python with a simple Flask dashboard.

The key design decision: the AI extracts and suggests, but the actual FCRA rules live in plain Python. No LLM guessing on legal classification. The agent's job is to do the reading, looking up, and calculating — the human still makes the final call.

We wired up five tools:

  • extract_receipt_data — pulls fields out of the receipt
  • lookup_travel_log — finds what the coordinator actually did that day
  • classify_fcra_expense — applies the split rules
  • check_admin_cap — keeps the 20% line safe
  • record_to_ledger — writes the audit trail

Challenges we ran into

Three things tripped us up.

First, Google deprecated the Gemini model we were using (gemini-2.5-flash) right in the middle of our build. Then gemini-3.6-flash hit a 20-request-per-day free tier limit that we burned through during testing.

Then xAI wanted a $5 minimum purchase to generate an API key. We pivoted to Groq, which is free, has no daily quota, and turned out to be faster than Gemini anyway.

Finally, the model ID we had for Groq (llama-3.3-70b-versatile) was retired on August 16, 2026. We switched to openai/gpt-oss-120b and then to llama-4-scout for cleaner output.

The harder design challenge was figuring out when the agent should stop and ask. We landed on three triggers: low confidence on the extraction, missing travel log data, and risk of breaching the cap. Anything else, the agent just handles it.

Accomplishments that we're proud of

  • It actually works. Full agent on Strands SDK, not a mockup.
  • The FCRA rules are deterministic Python, so the agent can't hallucinate on legal decisions.
  • The 20% cap check is real. The agent refuses to record breaches.
  • The pause is designed on purpose. It stops at exactly the right moments.
  • The ledger has a full audit trail — timestamp, split, and who confirmed.
  • Four test scenarios all work: mixed expense, pure programmatic, pure administrative, and HARD STOP.

What we learned

The biggest lesson: an AI agent shouldn't make legal decisions. The best thing we did was keep the FCRA rules in plain Python and let the AI only handle extraction and calculation.

The second lesson: the pause IS the product. The competition brief said agents should "surface only when there's a real decision to make." Designing when the agent stops turned out to be as important as designing what it does.

Third: simple wins. Five tools, one rules engine, one ledger. No vector database, no complicated orchestration. Everything was easy to debug and easy to demo.

Fourth: free-tier models are good enough. Groq handled every test we threw at it, with zero rate limits during our demo recording. For NGOs with tiny budgets, that matters.

Fifth: model IDs change fast. Always build with a fallback in mind.

Built With

Share this project:

Updates

Submission history