Inspiration
Spreadsheets are the most widely used programming environment on earth. They run budgets, forecasts, payroll, inventory, grant reporting and board packs. And they are the only programming environment with no tests, no linter, no code review and no CI.
The failure mode is always quiet. Someone inserts a row and a SUM stops
covering it. Someone types a number over a formula because it "looked wrong". One
cell in a column of three hundred drifted months ago. Nothing turns red. The
total still looks entirely plausible, and nobody finds out until the number is in
front of a board, a regulator or a funder.
This is not hypothetical:
- Public Health England lost roughly 16,000 COVID-19 cases because an XLS sheet hit a row limit and dropped the overflow silently. Contact tracing never reached those people.
- The Reinhart–Rogoff austerity paper's spreadsheet excluded rows from an average, and the conclusion shaped fiscal policy across several countries.
- JPMorgan's "London Whale" risk model contained a copy-paste error between two cells.
Everyone knows this. Nobody checks, because there has never been anything to check with.
The second half of the idea was realising it cannot be an app. Nobody would ever open a spreadsheet-linting tool — the defect is invisible by construction, so there is no moment at which a person thinks "I should go check." An app that waits to be opened will never be opened. The only form factor that can work is something that watches continuously and speaks once. That is exactly the hackathon's brief: runs autonomously, surfaces only when there's a real decision to make.
What it does
You point Plumb at a workbook. It watches.
It reads every formula, finds the ones that are provably odd, works out what each one is actually costing, and then decides that almost none of it is worth telling you about:
50,000 cells → 340 candidates → 8 defects → 2 surfaced
When it does speak, it leads with the consequence rather than the mechanism:
Q3 Budget!D11 — +270,734.00 SUM covers D2:D7 but the data continues to row 9 — 2 rows excluded
- =SUM(D2:D7)→+ =SUM(D2:D9)also changesSummary!B1
That figure is not an estimate. Plumb substitutes the corrected formula, re-evaluates the workbook, and measures the difference. "Possible error in D11" is easy to ignore; "your Q3 total is understated by 270,734" is not.
Because it keeps a history of every run, it can also say when: "the total stopped covering the last two departments on Tuesday." That points at a specific edit in a specific meeting. "The total is inconsistent" points at nothing.
Three ways to use it, and the first needs nothing installed:
| Who | How |
|---|---|
| Anyone with a spreadsheet | The web page — no install, no account, no upload |
| Someone who wants a file watched | plumb watch budget.xlsx |
| A team with workbooks in git | A GitHub Action that comments on the pull request |
How we built it
The central architectural decision, and the one everything else follows from:
Deterministic analysis finds candidates; agents judge the survivors.
Handing a whole workbook to a language model hallucinates, does not scale, and cannot prove anything. So Plumb does exact graph analysis in Python first, and spends model tokens only on the question that genuinely needs judgement.
Stage 1 — parse. The core primitive is A1 → R1C1 normalisation. In A1
notation =B2*C2 and =B3*C3 look like different strings, so a column of three
hundred such cells looks like three hundred unrelated formulas. Normalised to
relative notation both become =RC[-2]*RC[-1] — identical. Any cell whose
normalised form differs from its neighbours has drifted. That one transform
powers the highest-value detector in the suite.
Stage 2 — detect. Fourteen deterministic detectors: formula drift, hardcoded overrides, ranges that stopped covering their data, double-counted totals, numbers stored as text, missing formulas, error masking, orphaned blocks, cross-sheet drift, magnitude anomalies. No model, no credentials, ~4 ms.
Each declares how much its evidence is worth — proof, structural, or
distributional — and the weakest can never reach a human on its own. The
outlier detector uses median absolute deviation rather than standard deviation,
because a single catastrophic outlier inflates $\sigma$ enough to hide itself:
$$\text{score}(x) = \frac{|x - \tilde{x}|}{\operatorname{MAD}}, \qquad \operatorname{MAD} = \operatorname{median}\left(|x_i - \tilde{x}|\right)$$
Stage 3 — quantify. For every finding, substitute the correction, re-evaluate, diff. This needed a calculation engine, and the obvious choice was wrong (see below), so we wrote a narrow evaluator that answers exactly one counterfactual question.
Stage 4 — triage. Strands agents judge each candidate in context, returning structured output rather than prose. Context matters enormously here: a hardcoded number on a sheet of assumptions is how that sheet is supposed to work; the identical hardcode on a sheet of computed results means someone painted over a formula and the model stopped responding to its own inputs. An interpreter agent works out which kind of sheet it is looking at, then judge agents run concurrently, one per candidate.
Stage 5 — the gate, which is the actual product. Four filters — ruling, confidence, materiality, memory — take a run that raised hundreds of candidates down to the one or two things a person should see. It never repeats a finding. It never re-asks about something you dismissed. Every suppression is logged with its reason, so the silence is auditable rather than mysterious.
Writing back. Plumb can fix a formula, but writing into someone's financial model unattended is exactly the sort of thing that should never happen. Two independent guards, layered cheapest-first: Cedar authorisation, which is deterministic and evaluated first, and a HumanInTheLoop classifier behind it. A model's judgement is not a security boundary; it is a usability layer sitting behind one.
And then we deleted the install. The deterministic layer needs nothing but
openpyxl, so the whole thing runs in a browser under Pyodide. No server, no
upload, no account — and the privacy story gets stronger rather than weaker,
because the workbook never leaves the visitor's machine. A sceptic can verify
that by loading the page, turning off their wi-fi, and dropping the file. It
still works.
Challenges we ran into
Nearly every one of these was found by executing something rather than reasoning about it.
The clean corpus immediately caught three bugs the recall tests missed. We
built matched pairs — every defective workbook has a clean twin from the same
structure — so a finding on a twin is unambiguously noise. It found that
subtotals were being pulled into the blocks they summarise, and that the range
walker counted a total row as data its own SUM had missed. Both fired on
essentially every well-formed spreadsheet. The demo would have died on stage.
Excel is asymmetric about text, and that asymmetry is the whole danger.
="95600"*3.25 evaluates happily to 310700, while SUM over that same cell
contributes nothing. A column can therefore feed correct row-level products and a
wrong total simultaneously. We only found this by cross-validating our
evaluator against a reference engine, which disagreed by exactly $95600 \times
3.25$.
The obvious calculation engine was 40,000× too slow. The reference library took 29.7 seconds on a 43-cell workbook — unusable for something meant to watch a file. But Plumb never needs to recalculate a workbook; it needs to answer one counterfactual per finding. That is a far smaller problem, and a narrow evaluator does it in 0.7 ms, agreeing with the reference on 24/24 cells.
A denial of service, in a page we were about to make public.
=SUM(A1:XFD1048576) names Excel's entire grid — seventeen billion cells — and
we expanded ranges cell by cell. One uploaded workbook hung the process forever.
The robustness suite found it by hanging on its own first run.
A RecursionError that wasn't even an attack. A column of cascading
references (=A2, =A3, …) is an ordinary spreadsheet, and at 150 links it blew
the Python stack. Following one link costs a measured ten frames, so the cap is
now derived from the live recursion limit rather than guessed.
Cached values were paired with formulas by position. The two openpyxl load
modes disagree about a sheet's dimensions — read-only reported 100 rows where
formula mode reported 50 — so zipping them was correct only by accident. A
disagreement about which rows exist would have attached the wrong value to a
cell: invisible, and it would have corrupted every impact figure downstream.
Script injection in our own GitHub Action. The PR comment interpolated the
report into a shell command, and that report contains sheet names taken straight
from the workbook — which on a pull request is content a stranger chose. A sheet
named ev$(id)il was arbitrary code execution on the runner.
Telemetry that silently reported zero. We were reading token usage off the model stop response, which carries only the message and stop reason. It reported 0 tokens forever — worse than not measuring, because it looks like a working instrument. We would have shipped a claim about efficiency backed by a gauge stuck at zero.
Testing an agent layer with no model access. Rather than wait, we built a
scripted model that speaks the real Strands protocol — structured output arrives
as a tool spec named after the output model, answered with a toolUse block.
That let the whole path run offline and deterministically, and it caught the
telemetry bug within minutes.
Accomplishments that we're proud of
The numbers are measured and reproducible, not claimed. Over 400 matched pairs, regenerable from a seed and re-derived by CI on every push:
| Recall | 396/400 (99.0%) |
| False positives | 1 across 17,748 clean cells (0.006%) |
| Evaluator agreement | 24/24 vs a reference engine |
| Median scan | ~4 ms |
The false-positive rate is the headline number, not recall — and we treated it that way. A watcher that cries wolf gets muted, and a muted agent is a dead agent. When recall fell from 99.8% to 99.0% because a larger corpus revealed the weakest detector's true rate, we published the lower number and explained why, rather than quietly keeping the better one.
Every finding carries a real figure. Not a severity score, not a confidence percentage — the actual amount the number moves, computed by recalculating the workbook.
It needs nothing installed. The whole deterministic layer runs in a browser tab. Boot in ~1.5 s, analysis in under a second, and the file never leaves the machine.
The safety story is proven, not asserted. Reads run unattended, writes always stop and ask, a refusal leaves the file byte-identical, an approval backs up first — all verified end to end against a model double.
What we learned
Restraint is the product. Anyone can list everything suspicious in a spreadsheet. The engineering is in deciding almost none of it is worth an interruption — and in making every one of those suppressions auditable, so the silence is trustworthy rather than mysterious.
A claim about architecture has to be enforced, not written down. Our README
said the deterministic layer was independent of the agent layer. It wasn't:
pipeline.py imported agents at module scope and the package declared
strands-agents as a hard dependency. Porting to the browser was what exposed
it, because Pyodide simply refuses a native dependency. There is now a test that
fails if anyone re-breaks it.
You cannot find this class of bug by reading. The buffering bug, the silent telemetry, the injection, the recursion crash, the dimension mismatch — every one surfaced only when something was actually run, and several only when run against deliberately hostile input.
Token cost scales with anomalies, not cells — but that is an empirical claim, so we instrumented it rather than asserting it.
What's next for Plumb
An Excel add-in. This is where Plumb belongs: a task pane inside Excel, checking the workbook already open, no upload and no context switch. AppSource review takes weeks and the parsing layer needs a JavaScript sibling, so it did not fit six weeks — but it is the obvious next thing, and it is the version that reaches people who will never open a terminal.
Watching a shared drive. Real workbooks do not live on one laptop. They live on SharePoint, OneDrive or Google Drive, edited by five people. An agent that watches the finance team's shared folder and speaks up in Slack when a total breaks is the real product for an organisation — and it is where AgentCore Runtime's eight-hour sessions genuinely earn their place.
Deeper AgentCore integration. Memory behind the dismissal ledger so decisions outlive a session, Gateway for tool exposure, and Policy enforcing the Cedar rules at the gateway rather than in-process.
More detectors, and evidence for each. Every new one has to prove itself against the clean corpus before it ships. That bar is why the false-positive rate is where it is, and it is not moving.
Built With
- amazon-bedrock
- amazon-bedrock-agentcore
- amazon-web-services
- cedar
- claude
- css
- docker
- fastapi
- github
- github-actions
- html
- javascript
- micropip
- openpyxl
- pillow
- pydantic
- pyodide
- python
- ruff
- strands-agents
- uvicorn
- webassembly
Log in or sign up for Devpost to join the conversation.