Inspiration
I run a Thai agri-export company that invoices Chinese buyers in USD/CNY. Every month a deposit lands with the wrong amount and a useless reference — "INWARD REMIT REF 88213" — and I have to guess which invoice it was, work out how much the bank silently skimmed, and remember which client always pays like this. Freelancers Union puts late/non-payment at 71% among cross-border earners; agri exporters live the same problem, just with bigger numbers. Nobody builds AR tooling for this — Xero/QuickBooks assume clean domestic deposits.
What it does
GotPaid ingests a folder of sent invoices (PDF/text) and a bank statement, then:
- Extracts structured invoice fields with Claude (the only job the LLM has)
- Reconciles deposits to invoices with a deterministic matcher tolerant of shorted fees, FX conversion, batched payments, and partial payments
- Learns per-client payment habits over time (typical fee bite, days late) — matching gets measurably better as history accumulates
- Outputs an evidence-backed unpaid list, a "hidden fees lost this year" total, and an auto-drafted chaser letter per unpaid invoice
How we built it
Python 3, stdlib only, Anthropic API via curl — zero pip dependencies. Deterministic synthetic data generator (10 fictional exporters, planted hard cases: batched payments, partial payments, an ambiguous same-amount "twin" invoice, duplicate wires, unpaid invoices) so every claim in this writeup is reproducible from logs. Foxit powers the closing move: the PDF Services API turns a drafted chaser letter into a real PDF, and the eSign API creates a draft signature folder — "agent starts from a prompt, ends with a document ready to sign," with a human reviewing before it's actually sent.
Challenges we ran into
- Our first matcher was a straightforward greedy sort — it looked fine on small companies but broke on our largest one (152 invoices): near-identical amounts from different buyers kept "stealing" the correct invoice away from a deposit that scored fractionally higher elsewhere in the sort order. Replaced it with a proper stable-matching pass (Gale–Shapley), which fixed the mis-assignments.
- Our amount-scoring formula had a subtle bias: it scored candidates by distance from the middle of a plausible bank-fee range, but real fees cluster near the low end, not the middle — so a coincidental match that just happened to sit near the midpoint kept beating the true one. Recalibrating against our own data's actual fee distribution fixed it.
- We built a single-prompt LLM baseline to compare against (paste everything, ask Claude to match it all in one call). On 3 of 10 companies it returned zero answers — the same "reasoning exhausts the output budget before writing anything" failure I'd already hit in an earlier hackathon project on a completely unrelated task. Seeing it twice is strong evidence multi-pass beats single-shot here.
- Foxit's eSign API is documented as needing a separate, sales-gated account. We planned a PDF-only fallback around that — it turned out to be instant self-serve for this hackathon's sandbox, so the fallback wasn't needed after all.
Accomplishments that we're proud of
- The full pipeline — extraction, reconciliation, two baselines, an honest 3-way eval, a dashboard, and real Foxit PDF+eSign integration — built and validated end-to-end in one sitting.
- Our engine beats both a no-AI baseline and a single-LLM-prompt baseline on F1 (0.738 vs 0.401 vs 0.562) and, more importantly, on money actually reconciled correctly (65% vs 22% vs 39%).
- Every number above regenerates from the repo —
python3 eval.pyscores all three approaches against a held-out ground truth none of them ever see.
What we learned
- A scoring heuristic that "sounds reasonable" can be quietly wrong in a way that only shows up at scale — we didn't catch the fee-distribution bias until testing on our largest company. Ground every threshold in the real data, not intuition.
- Single-shot LLM prompts don't just get less accurate on hard tasks — they can silently return nothing once reasoning eats the token budget. That's now a pattern we've seen across two unrelated projects.
- A sponsor's real sandbox can behave differently from its public docs (eSign wasn't sales-gated for us, despite the docs saying so) — always verify against the live account instead of trusting documentation when they conflict.
What's next
This is a standalone engine today; it's built to plug into LaungPro, our deployed durian-export platform (2 awards at the Nanning ASEAN summit), which already tracks invoices but has no settlement reconciliation yet.
Log in or sign up for Devpost to join the conversation.