Reckon

(Track 2: Autonomous Office of the CFO)

The number that started it

The US Government Accountability Office puts improper payments across 64 federal programs at $186 billion for FY2025. The UK National Audit Office puts its own figure at up to £81 billion for 2023-24. India's Comptroller and Auditor General published an audit paragraph in 2025 whose title is, in full, "Double payment to contractors for an item of work due to incorrect estimates and measurements."

So I went and read what is sold to fix this. Supervizor, AppZen, Oversight, MindBridge. Every one of them leads with the same sentence: we test 100% of your transactions. And every one of them does. That problem was solved a decade ago.

Here is what none of them publish anywhere on their site: a false positive rate.

I found the number in the academic literature instead. Published precision for continuous auditing exception detection is 0.73, at recall 0.18. The authors say plainly that most exceptions surfaced at that precision are false alarms.

Then I found the study that made me actually build this. Li, Brazel and Gold, Foundation for Auditing Research, 2024. 125 auditors, controlled design. The ones using full-population testing acted on a fraud red flag 48% of the time. The ones using sampling acted on it 65% of the time.

Testing everything made them less skeptical. Read that again. The tool designed to catch more caused people to chase less.

That is the whole problem in one number. Detection is not the bottleneck. The queue nobody works is the bottleneck. If your system produces 1,268 alerts and your finance team is four people, the alerts sit there and the money stays gone.

What Reckon does

Reckon is an accounts payable exception agent, and its actual job is to stay quiet.

Seven stages. Six of them are plain Python. Twenty thousand ledger rows collapse into 3,425 real payments, because one invoice can be dozens of distribution lines. Hash-bucketed blocking finds 1,279 candidate pairs. Deterministic rules kill 11 of those outright, the ones you can prove innocent with a lookup. That leaves 1,268 candidates worth $649,739.73.

Then, and only then, a model looks at what is left.

Its instruction is inverted from every tool in this category. It is not asked to find duplicates. It is asked to find the innocent explanation. A recurring monthly invoice for the same amount. Two purchase orders differing by one character. A payment whose parts add up exactly to the invoice total. A description reading "Progress payment 2 of 4", which is a scheduled construction draw and is supposed to happen four times.

On the committed sample it dismissed 388 of them. Dollars at risk fell to $170,157.90. 800 escalated, 5 errored, and 75 had no matching recording and were surfaced rather than quietly waved through.

Every survivor goes to a named person in the department that owns the payment, with the evidence chain attached. Nothing posts automatically, ever. Confirmed findings produce a claim letter, a journal entry and an audit trail.

The letter is my favourite thing in the repository:

This is a query, not a claim that an amount is owed. You may have a perfectly good explanation, and we would rather ask than assume.

A tool built to maximise clawbacks would never write that sentence.

Track: Autonomous Office of the CFO (Track 2). Reckon automates the accounts payable payment review that runs before each payment run, and the monthly recovery sweep. Exception handling and human review are the product here rather than a feature bolted on at the end.

The data is real and I read every column

Two published government ledgers, both verified by downloading them and reading the actual schema rather than trusting a description of it.

Checkbook L.A., published by the City of Los Angeles Controller under CC BY 4.0. 6.47 million rows. Twenty thousand of them are committed in the repository so anyone can run the whole thing. It is the only public payment ledger I found that carries a real invoice number, a payment status, a purchase order reference and a publisher-assigned vendor_id all at once. That last field matters more than the others, and I will come back to it.

Oklahoma's state vendor payments file is in there specifically because it is awkward. No invoice numbers at all. cp1252 encoding. Roughly a quarter of the vendor names redacted by the publisher. Each adapter declares what its source lacks, and blocking disables the signals the data cannot support. Portability is an interface, not a sentence in a README.

I went looking for Indian data properly, and mostly it does not exist in this form. Seven central sources, all captcha-gated, login-walled or aggregate only. Bengaluru's municipal corporation does publish a genuine contractor payment ledger, around half a million lines. All of that is written up in docs/DATA-SOURCES.md, including everything that did not work, because the failures are the useful part for whoever tries this next.

How I measured it, and the number I am not hiding

Claims are cheap. I planted 40 real duplicates and 60 decoys into the ledger.

Every decoy is deliberately built so that no deterministic rule can resolve it. A cancelled payment. A progress draw. Two POs differing by one digit. A recurring invoice whose only tell is monthly cadence.

The baseline is the identical pipeline with stage 5 removed. Same aggregation, same normalisation, same blocking, same deterministic dismissal. The only variable is whether an agent adjudicates the residual.

metric rules only full pipeline
precision 0.400 0.718
recall 1.000 0.700
F1 0.571 0.709
dollars wrongly flagged $76,093.78 $0

The rules engine fell for 60 decoys out of 60. Reckon fell for 11. Cancelled payments, line splits and progress payments went to 0 out of 10 each.

Now the number that goes the other way. Recall fell from 1.000 to 0.700. It misses three real duplicates in every ten. That is what the precision costs, and it sits in the README in the same size type as the wins. A tool that flags everything has perfect recall. That is the tool nobody uses.

One more, and this one is not graded by a model at all. The LA ledger carries that publisher-assigned vendor_id, so vendor name canonicalisation is scored against the city's own labels rather than my opinion. Precision 0.967, recall 1.000, F1 0.983. Real chaos it has to survive: W W GRAINGER INC, W. W. GRAINGER INC., W. W. GRAINGER, INC. and W.W. GRAINGER INC. are one vendor. 150 collision groups covering 319 spellings in a 200k sample.

How I built it

Entirely with AO, Agent Orchestrator. 35 sessions, 26 merged units of work, 347 tests. Each worker ran in its own git worktree on its own branch, and every one of those branches is pushed to the repository, so the build process is verifiable from the repository alone.

The thing that actually made parallel agents work got written before any code. I wrote docs/STANDARDS.md first, an engineering contract that every worker reads as part of its brief. Complexity limits. Bounded memory. Decimal for money, and a validator that rejects a float on a money field outright. Typed errors. mypy strict. Performance budgets expressed as tests rather than as prose.

Then docs/SPEC.md, which locks the data contract. Transaction, Candidate, Verdict, ReviewDecision, GroundTruth. Two workers writing to the same schema in parallel produce code that composes. Two workers inventing their own schemas produce a merge conflict you spend a night untangling.

Late in the build I stopped spawning workers by hand and gave the orchestrator an outcome instead of a task list: make the demo output self-verifying. It planned the work, spawned four workers, and all four shipped.

Some engineering worth naming.

All-pairs comparison over 700,000 transactions is 2.4 × 10^11 comparisons, which is why nobody does it. Blocking is hash-bucketed by composite key with a hard cap per bucket, and splits by a secondary key when a single vendor grows too large. There is a test that generates 200,000 rows and asserts the whole thing finishes in under 30 seconds.

Adjudication is two-tier. A fast model handles the volume, and anything below 0.7 confidence escalates to a stronger one. Calls run through asyncio.gather behind a semaphore.

Every model call made during development was recorded, keyed by a hash of the request payload. That is why you can clone this repository and run the entire demo with no API key at all, offline and deterministic. It is also why I could diagnose the worst bug of the build, which is coming up.

Nothing is hardwired to one vendor. RECKON_API_BASE, RECKON_API_KEY, RECKON_MODEL_FAST and RECKON_MODEL_ESCALATE are the entire configuration surface, so it points at whatever endpoint you already pay for.

What broke

Three concurrent workers deadlock on this machine. Two never do. I lost 7.6 hours to three workers burning CPU and producing exactly zero files before I worked that out.

Workers sometimes finish the work and then stall before committing it. The code was sitting complete in the worktree, so I committed it on the worker's behalf.

Chat mode cannot launch an agent whose CLI is a .cmd shim on Windows. It dies about a second after spawn. TUI mode goes through ConPTY and works.

Then my favourite bug. 152 verdicts came back errored. It read exactly like model unreliability, and I nearly wrote that into the README as a limitation. It was something else. One field was returning a nested list where the schema expected a flat one, and strict validation refused to guess at it. The model had been right every single time. Fixing the coercion took the error rate from 8.8% to 0.17%.

And the mistake I made rather than inherited. I wrote a diagnostic script to check cassette matching, and it built its request payloads slightly differently from the real adjudicator. It reported zero matches. On the strength of that I reverted a prompt change that was actually good. I caught it by opening a cassette file and reading it by hand. A diagnostic that does not share code with the thing it diagnoses is a second implementation you now have to trust, and I had not earned that trust.

The bug that scared me most surfaced last. Replay silently failed when no environment variables were set, because the model name is part of the payload hash and the default was a placeholder. Every local test passed. A judge cloning the repository would have watched every candidate error out. I only found it by cloning fresh from GitHub with a stripped environment, which I then did three times before I believed it.

What I learned

Precision is the product. Recall is a slider, and anyone can push it to 1.000 by flagging everything. The hard problem was always the part where you decide to say nothing.

Strict validation surfaces bugs that lenient parsing hides. If those 152 verdicts had been silently coerced I would have shipped a slightly wrong number and never known it was wrong.

Parallel agents need a written contract far more than they need clever prompting. The standards document did more for output quality than any instruction I wrote into an individual worker brief.

And test the clean-clone path early. Everything working on your own machine is not evidence of anything.

Honest limits

Recall is 0.700. Recurring-cadence decoys are the weakest case at 6 misses out of 10, and that is the next thing I would fix.

I measured a prompt improvement that raised escalations from 19 to 40 on the 50 hardest cases, but adopting it meant re-recording 1,731 model calls and the clock ran out. It is documented in the README rather than quietly dropped.

Everything this produces is a candidate for human review. It makes no claim that anyone did anything wrong. Public ledgers contain payments to individuals, so names are masked.

Every number above has a command in the README that reproduces it.

github.com/tanayvasishtha/Reckon

Built With

Share this project:

Updates

Submission history