A label or a distributor sends a payee a list of tracks, a play count for each, a rate and a total. Almost nobody reconciles that document against the usage behind it, because reconciling means running millions of play records against a statement that was never designed to be audited. So the statement gets trusted by default. Against a 24 hour window of 1,174,697 plays carrying 30 defects planted by a script it can never read, recoup found 30 of 30, with no false positives and the evidence matching the plant every time. Those 30 defects account for $103.60 of a $1,180.47 statement, 8.78% of the money paid, under the declared demo rate card.

recoup is an agent that does the asking. It holds a usage ledger and a royalty statement in ClickHouse Cloud, runs five detectors against them, and emits every discrepancy together with the SQL query that proves it. Four of those detectors look for money that does not add up: usage carrying no canonical identifier, lines claiming fewer plays than the ledger supports, lines claiming usage the ledger does not contain at all, and lines where plays multiplied by the declared rate does not equal the amount paid. The fifth detector is the one this project is actually about, and it does not look for errors at all. It reports how much of the statement cannot be verified in the first place.

Being honest about the data matters more here than the code does. The listens are real. They come from a public ListenBrainz incremental dump, released CC0 by MetaBrainz, and recoup declares that corpus the usage ledger of record for the period. That is a modelling decision stated out loud, not a claim that ListenBrainz mirrors what any streaming service actually paid on. It does not. The statements are generated from that ledger by a script that injects defects deliberately and records every one in a ground truth file the agent can never read. No real listener, artist or service is accused of anything anywhere in the system. Findings are about statement lines and injected defects. The test fixtures are real records with the user identifiers replaced by stable derived tokens, because CC0 makes republishing legal but the tests never needed that much personal data.

The refusal is the interesting part. Payability is not a yes or a no. It has four values: payable, not payable, unmeasurable, and incoherent. A row is unmeasurable when the rate card asks for a field that row does not carry, and the refusal names the missing field rather than guessing at a value. A row is incoherent when its played time exceeds its own track length, which is not a large payment but a contradiction, so recoup reports it and never quietly clamps it. The same discipline runs through the whole system. A query that failed is recorded as unmeasured, never as zero rows, because a read that did not happen proves nothing about the rows it did not reach.

The measured result, with its population named, is this. Across the 24 hour statement window of 13 August 2026, 1,173,427 of 1,174,697 plays, or 99.89%, are unmeasurable under the declared demo rate card, which pays on 30 seconds of played time. That number never travels without its diagnosis, because on its own it reads as though the data were broken, and it is not. Coverage turns out to be a property of the reporting client rather than of the period. The clients submitting into that particular window do not supply played time, while one long running importer supplies it on between 97.66% and 100% of its rows across all 11 years it spans. Coverage predicted from each year's client mix tracks the actual figure to within 0.58 percentage points. A different figure describes the whole corpus rather than the window: 53.59% of all 3,820,939 listens, whose timestamps span 2005 to 2026, carry neither a track length nor a played time at all. Those two numbers answer different questions over different populations, and they are never quoted together.

Against the injected defects the detectors found 30 of 30, with no false positives and with the evidence matching the plant on every one. That result needs its caveat attached, because the caveat is load bearing. The generator injects defects using the same rules module the detectors use to find them. So a clean sweep proves the pipeline works from end to end, and does not prove the rules are correct. A shared rule that is wrong in the same direction on both sides scores 100% while being wrong, and the scoreboard cannot see that error by construction. The evidence that is not circular is the measurement of the real data, where nothing was injected: 90.95% of the corpus carries neither a recording identifier nor an ISRC, and the client mix finding holds across 3.8 million rows that nobody planted.

The architecture exists to make one claim defensible. The model has no code path to a verdict. Each detector is split in two, so pure code builds the SQL before anything runs and pure code decides every verdict from the returned rows alone. The agent, running on Vertex through the Google Agent Development Kit, sits between those halves and does nothing but carry queries to ClickHouse through the official mcp-clickhouse server and carry the rows back. Its tools take no arguments at all, so it cannot write a query, pass a value into one, or name an outcome. It connects as a database user granted nothing but SELECT, and the proof of that is adversarial rather than trusting: 18 attempts to write, across every configuration tried including one with the write permissions deliberately forced on, all failed, while the same statements run as an administrator succeeded 9 times out of 9.

One last thing was proved by switching something off. The hosted report page serves a recorded campaign from a committed snapshot and carries the timestamp of when that campaign was collected. The container ships with no database driver and no credentials, so it could not query the ledger even if a later change asked it to. On 23 August the ClickHouse service was stopped deliberately, the page was fetched while it was down, and every content check passed. When the database is gone the report survives it, and that was demonstrated rather than assumed.

Built With

Share this project:

Updates

Submission history