What it does

Stratra is an accounts-payable agent that starts with no rules configured at all.

It reads a company's own artifacts — twelve months of journal entries, the close memos, the Slack threads where decisions actually got made — and reconstructs the accounting policy as an explicit, versioned ruleset. Every rule carries the evidence it was derived from. It then uses that ruleset to code new invoices, posting the ones it can justify with a citation and escalating the ones it can't to a named person. Each human answer becomes a new rule, so the same question is never asked twice.

Two workflows are live: invoice GL coding and accrual cutoff.

Inspiration

An ERP records that an invoice was coded to 6100. It has never recorded why.

That reasoning — why this vendor is COGS and that one is opex, why the treatment changed in March, what counts as material at 11pm on day four of the close — lives in a controller's head, in an inbox, and in a file called FINAL_v7_USE_THIS. Maximor calls it the human runtime, and their framing is the thing we built against: the system of record captures the outcome of a judgment and nothing of the judgment itself.

The second input was the trust gap. 96% of CFOs want AI freeing their teams for strategic work; only 14% trust it to deliver accurate accounting data on its own. That gap is not about model intelligence. It is about whether a system knows the boundary of its own confidence.

So we didn't build a system that codes invoices well. We built one that refuses to code an invoice it can't justify, and made that refusal the product.

How we built it

history + workpapers + Slack
            │
         [miner]      reconstruct policy, flag self-contradiction
            │
        [ruleset]     versioned, every rule carries its evidence
            │
        [executor]    most-specific rule wins → post with justification
            │
       [escalator]    NO_RULE | CONFLICTED | AMBIGUOUS |
                      LOW_CONFIDENCE | MATERIALITY | PERIOD_AMBIGUOUS
            │
    human answer → [fold-back] → new version, never asked again

History mining is deterministic; document and correspondence mining use a model. Counting that 49 of 49 entries posted to one account is something code does better than an LLM. Reading a prose close memo and spotting that it contradicts a memo from three months earlier is not. We kept the deterministic path as a baseline and report both.

Inference runs on Groq (openai/gpt-oss-120b) behind a provider-agnostic layer. Mining a full corpus costs $0.00144 and 6.2 seconds, measured. Execution is rule matching with no per-invoice model call, so cost does not scale with invoice volume.

Everything from the first commit onward was built through six AO sessions in isolated worktrees: the live re-run server, the regression suite, the accrual workflow, the cost instrumentation, and the docs.

Three design decisions worth defending

It refuses to average away a contradiction. When a vendor's history splits, the miner tries to explain the split by entity, then by date. If neither explains it, it emits a CONFLICTED rule and escalates. Taking the majority treatment would be right ~90% of the time — but the 10% is the capitalisation decision, and that is the one that becomes an audit finding.

It never generalises past its evidence. A rule learned from a human answer may only retire a rule that is no more specific than itself.

A learned threshold is an interval, not a point. When answers separate on amount, the agent asserts only what it has confirmed — at or below 6,117, at or above 27,974 — and the untested band between keeps escalating until someone narrows it. The true threshold is 10,000 and the band correctly brackets it. It does not pretend to know where the line is.

Results

Posted without asking 152 / 200
Account coding accuracy 100%
Accrual accuracy 152/152 booked to the right month
False autonomy (posted and wrong) 0
Rules reconstructed from zero 20, across 11 versions
Mining cost / time $0.00144 · 6.2s (Groq, measured)

Across 22 independently generated corpora — 12 of which were never used during development — false autonomy is 0.

Challenges we ran into

The single-seed run lied to us. Our first run reported zero false autonomy and we nearly shipped that number. Then we wrote sweep.py to regenerate the entire corpus across many seeds — and found 49 wrongly-coded invoices.

All three root causes were the same species: generalising past the evidence.

  1. One answer about a large licence became a vendor-wide rule and mis-coded every small subscription.
  2. An answer at 6,467 became "everything above 6,467," swallowing 25k licences. The fix is deliberately asymmetric — the low side is bounded at zero so amount_max is safe, but the high side has no ceiling, so a lone answer there generalises to nothing until an answer on the other side brackets it.
  3. Two answers landing in different entities let the system "explain" an amount threshold as an entity difference.

We also found that a 95% "clean rule" threshold silently absorbed a 3% minority as noise — and the rare exceptions are exactly the material ones. It is now zero-exceptions-or-escalate.

The cost: escalations rose from ~15 to ~18 human answers per corpus. That is the correct trade.

Adding accrual cutoff dropped autonomy from 179/200 to 152/200, because period cutoff on a boundary-spanning invoice is a judgment we won't fake — and unlike account rules, it doesn't generalise, so it doesn't decay across batches. We kept it and report it rather than quietly demoing the higher number.

What we learned

Escalation rate is a cost metric. False autonomy is a trust metric. Only one of them has to be zero, and it is not the one that makes a demo look impressive.

We also learned that an eval which runs once tells you almost nothing. The sweep was the single highest-value thing we built — it found bugs that a working demo would have hidden all the way through judging.

What's next

Cost-centre assignment, intercompany eliminations, and audit PBC fulfilment — all processes where the policy was never written down. The corpus here is synthetic with three landmines planted deliberately; real posting history is messier, and the next step is a run against a real ledger export.

Built With

Share this project:

Updates

Submission history