Inspiration

Month end close is five to ten days of a controller doing the same three things they did last month. Tying the bank statement to the ledger. Deciding which account and department each vendor invoice hits. Hunting for costs that were incurred but never invoiced. The judgment is real, but it is also identical every month, and it lives in one person's head. Rules engines in NetSuite and QuickBooks are too brittle to hold it, so nobody writes it down.

What it does

Tickmark takes over that close end to end. It codes AP invoices, reconciles the bank across all four match shapes, and finds the recurring costs that never got billed. Anything it cannot assert goes to a controller in an exception queue. The part we care about most is what happens next. When a controller resolves an exception and says always do this, Tickmark does not create a rule. It answers: noted, but not yet a rule, one correction is an anecdote, I will propose a rule the next time a correction agrees with it. The second agreeing correction proposes the rule, citing both verbatim, already replayed against every closed period. Adoption names its approver, and a rule that failed its replay cannot be adopted at all.

That matters because most agents learn by growing a prompt, so every lesson makes the next call more expensive and none of it can be audited. Tickmark compiles the judgment into a predicate and an action that executes with no model call. Learning makes the close cheaper instead of dearer. It covers the six workflows the track names. Closing the books, reconciling accounts, processing invoices, gathering audit support, preparing cash reports, and updating forecasts. All of them run on books you import yourself.

How I built it

I ran the build through AO from the start. An orchestrator session planned the work and delegated it to worker sessions in isolated git worktrees, most of them running Codex, with Claude Code sessions working alongside them. The stack is Next.js 16 and React 19 on Vercel, Supabase Postgres as the system of record, Clerk for auth, and the AI SDK talking to Claude. TensorMux is the inference path and Neatlogs traces the agent runs.

The accounting invariants are enforced as database triggers rather than prompt instructions. Debits must equal credits. Tickmarks are append only, so you supersede one rather than edit it. Preparer cannot equal approver. No rule activates without a backtest and evidence from at least two distinct corrections. I integrated Dodo Payments through their payout breakup export so processor settlements reconcile properly. A processor pays out net of its fees, so the deposit never equals the revenue behind it. Each payout expands into the three rows that actually explain it. Bank at net, revenue at gross, fee as an expense.

Challenges I ran into

Almost every real bug came from running the thing rather than reading it.

A rule learned in March fired zero times in April, because Gusto appears as GUSTO PAYROLL one month and GUSTO TAX COLLECTION the next. A predicate cut from one statement's wording does not survive to the next period, which is the exact failure this product exists to prevent. The importer silently read an unquoted 1,240.00 as one dollar, because the thousands separator split the CSV field on itself. Reading twelve hundred as twelve is an accounting error, not a parsing inconvenience, so a row whose field count disagrees with its header is now refused with that reason. The accrual feature was structurally dead. Recurring history was built from the current period's invoices, so every vendor in it had by definition already billed and nothing could ever be found missing.

Accomplishments that I am proud of

Everything is measured rather than asserted. Across four closes on the same books with the same model, auto cleared went from 55 percent to 79 percent, exceptions fell from 41 to 21, and auto clear precision held at 100 percent. That last number is the one that must not move, because nobody reviews what the agent cleared unattended.

There are eight verification suites and one command that runs them all. They prove the guarantees by deliberately violating each one and requiring the database to reject it, and they prove the whole learning loop turns on ingested books with no model involved at all.

What I learned

Refusing to act is a feature. The strongest moment in the product is the one where it declines to learn from a single example and explains why. Accountants trust that far more than confidence.

What is next for Tickmark

Multi entity support, a live bank feed instead of CSV import, and intercompany eliminations.

How I used AO

I ran the build through AO from the start. The repository was registered as an AO project and an orchestrator session planned and delegated the work, spawning worker sessions into isolated git worktrees so several agents could work the same codebase without stepping on each other. Across the hackathon that came to 20 worker sessions and 3 orchestrator sessions, each on its own branch under ao/tickmark-N. Most of the workers ran Codex, and 13 commits in the repository landed through them, covering the verified measurement report, structured output support, evidence rules for unattended close, rule deduplication, and the model outage handling. The remainder ran in Claude Code sessions alongside them.

AO doctor confirmed the harness setup before we started, both Codex and Claude Code were detected, and the daemon ran throughout. The kanban view is what we used to track which branches were iterating, which were in review, and which were ready to merge.

Built With

Share this project:

Updates

Submission history