We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Try it in 30 seconds: open civicprobe.vercel.app, click a salary under Try, then type 481,000 into "What your calculator said". Code: github.com/Abdullah49645/civicprobe

Track: ⚖️ Access to Justice & Civic Tech (also relevant to 📜 Digital Rights & Policy Tech)

CivicProbe taxpayer check

Inspiration

Every year, governments change tax rates, benefit thresholds and eligibility rules. The law changes on a set date, but the websites and calculators people use to follow it are updated by hand, by whoever maintains them, whenever they get around to it. Nobody checks that they got it right.

Pakistan's Finance Act 2026 is a good example. For tax year 2026-27 it changed the rates for salaried people, added two new slabs, and removed a 9% surcharge on higher incomes. Millions of people use online calculators to estimate their tax, plan budgets or check their payslip. A calculator that updates the rates but forgets one slab, or still applies the old surcharge, gives confidently wrong answers, and nothing on the page tells you. While we were building this, some public calculator pages still described the surcharge that had just been removed.

This isn't a Pakistani problem. It happens wherever rules change and software has to catch up. We wanted a way to check it.

The problem and our solution

Problem: when a law changes, there is no easy way for an ordinary person to know whether the calculator in front of them follows the new law, and no systematic way for anyone to audit the software that citizens rely on.

Solution: CivicProbe compares the old and new versions of a rule, works out exactly which answers the change should affect, and checks the software at precisely those points. It works at two levels:

  • For taxpayers: type in your salary and what your calculator told you, and see whether it matches the new law, the old one, or neither.
  • For auditors: point it at a calculator and it runs a full, reproducible audit that pinpoints which part of the new law the software missed.

What it does

The taxpayer check is the first thing on the site, in English or Urdu. Enter your salary (yearly or monthly) and you see your 2026-27 tax, which slab you fall in, the formula behind it, and how it compares with last year. If you also enter what your calculator said, CivicProbe tells you whether it is right, still using last year's rules, or simply wrong. It can also draft a polite note to the calculator's owner. Everything runs in your browser; nothing you type is sent anywhere.

For example, on a salary of Rs. 3,650,000 the new law gives Rs. 428,500. A calculator still on last year's rules says Rs. 481,000, which is Rs. 52,500 too much. CivicProbe recognizes that exact figure as last year's law and says so.

The audit engine does the same thing systematically:

  1. Finds exactly what changed. It reads both versions of the tax table and works out the precise set of incomes whose tax must now differ, $$\lbrace\, x : T_{\text{new}}(x) \neq T_{\text{old}}(x) \,\rbrace$$ using exact fractions rather than floating point, and respecting the statute's wording ("exceeds X but does not exceed Y"). For the Pakistani table, it finds 11 separate edits.
  2. Chooses where to test. Instead of random salaries, it picks the ones most likely to expose a mistake: where old and new differ most, right around each changed threshold, and inside every changed slab. Each test records why it was chosen.
  3. Runs the tests in a real browser, filling in the calculator's form like a person would. The browser is read-only: it can't navigate to other sites, it respects rate limits and robots.txt, and it never submits anything that changes data.
  4. Shrinks any failure to the smallest salary that shows it, so a person can check it by hand, and states exactly what it proved.
  5. Explains the failure. It builds a list of possible single mistakes (one old rate left in, one slab never added, the old surcharge still applied, and so on) and keeps only the ones that reproduce every answer the calculator gave.

In our main demo, the calculator had one slab's rate left at 30% instead of 25%. CivicProbe caught it on its second test, shrank it to Rs. 3,200,025 (where the error first shows), found that the same mistake reaches Rs. 45,000 higher up, and out of 58 possible explanations identified the one that fits all 81 answers: that stale rate.

Audit verdict and diagnosis

Other things reviewers can try on the live site:

  • Replay: one click re-runs a recorded audit inside your browser and confirms every result matches the original run, so you don't have to take our word for it.
  • Test a calculator thoroughly: type any real calculator's answers for 19 chosen salaries, and CivicProbe diagnoses it on the spot.
  • Evidence for every result: each test links back through the rule, the change it targets, the expected and observed values, and the source it came from.

It also handles a calculator that hasn't added the new tax year at all (reported as "not updated", with no tests sent), and a second, non-tax rule type, eligibility based on age and income, to show the engine isn't tied to tax.

How we built it

  • TypeScript, split into small packages: exact arithmetic, policy engine, change ("delta") engine, test planner and shrinker, comparator, diagnosis, and an evidence format.
  • Playwright with Chromium drives the calculators through a configurable form adapter, the same one a live site would use.
  • The law as data. Both versions of the tax table are encoded with their sources and checked against all 9 worked examples in KPMG's Brief of Finance Act, 2026, plus 5 independent examples from PakFiler.
  • The website is a single static page built with esbuild and deployed on Vercel. The whole engine runs in the visitor's browser, with no server and no accounts.
  • Testing: 56 unit and property tests (including 200 randomly generated tax tables checked against brute force) and 21 end-to-end tests in a real browser. CI checks that the website build is byte-for-byte reproducible.

To check the method actually helps, we built a benchmark. We generated 60 subtly broken calculators automatically from the two versions of the law. Three of them turn out to be impossible to detect after rounding, so we excluded those. With 50 tests each, CivicProbe caught all 47 remaining faults; random testing caught 36.7% and classic boundary testing 34%. We also report the cases where random testing does better.

Benchmark results

Challenges we ran into

Rounding breaks the obvious approach. Calculators round to whole rupees, so a wrong rate can show an error at one salary and none a few rupees higher. That means the usual "cut the range in half" search can return the wrong smallest case. We saw it in our own results: Rs. 3,200,025 fails while Rs. 3,200,029 passes. We wrote a shrinker that walks the law's own thresholds first, then checks every remaining value, and reports exactly what it verified.

Testing right at a threshold isn't enough. Our first planner missed a calculator with a typo in one slab boundary, because right next to a threshold the error is smaller than a rupee and rounds away. If the marginal rate changes by \( \Delta r \) at a threshold and the comparison allows a tolerance of \( \tau \) rupees, the error only becomes visible after $$k = \left\lfloor \frac{\tau + 1}{\Delta r} \right\rfloor + 1$$ rupees. The planner now places tests at that distance on both sides of every changed threshold, and the benchmark checks typos at every threshold, not just the one that caught us out.

Staying honest. It would have been easy to present the demo results as real-world findings. They aren't: the calculators in the recorded audits are practice versions with deliberate mistakes, and they're labelled that way everywhere. The encoded law is marked "corroborated" rather than final, because we checked it against published examples but haven't hashed the official gazette text.

Accomplishments that we're proud of

  • An ordinary taxpayer can catch an out-of-date calculator in about ten seconds, in English or Urdu.
  • CivicProbe says which part of the new law a calculator missed, not just that it's wrong.
  • Every result on the site can be recomputed in the reviewer's own browser.
  • The benchmark shows where the approach works and where it doesn't.

What we learned

Legal text is more precise than it looks. The difference between "exceeds" and "does not exceed" decides which slab a salary falls into, and ordinary software testing tends to ignore that. The most useful idea turned out to be simple: don't test everything, test the difference between the old law and the new one. That turns a vague question ("is this calculator right?") into a specific one that can actually be answered.

We also learned how much trust depends on being able to check the work. Replay and the evidence chain changed how the whole project felt.

What's next for CivicProbe: The law changed. Did the software?

  • Beyond Pakistan: the engine works with any progressive tax table or eligibility rule. Next up: UK income tax bands, US state tax tables and benefit eligibility screeners.
  • Real calculators: run the read-only live runner against public Pakistani tax calculators, and report anything found privately to their owners first.
  • Lock the law: tie the encoded table to the official published text of the Act.
  • After every budget: work with civic groups and legal clinics to check government e-services and popular calculators whenever the rules change.

Transparency

  • Libraries: Playwright (Apache-2.0), TypeScript (Apache-2.0), esbuild (MIT), tsx (MIT). The arithmetic and hashing are written from scratch.
  • AI tools: we used Anthropic's Claude as a coding assistant during development. No AI model runs in CivicProbe itself. Every result comes from deterministic code and can be reproduced from the published evidence, which matters when the question is "what does the law say?".
  • Data sources: the Finance Act 2026 salaried tax table, cross-checked against worked examples published by KPMG Taseer Hadi & Co. and PakFiler.
  • Demo calculators: the calculators in the recorded audits are practice versions with deliberate mistakes, clearly labelled. No real service is accused of anything.

Built With

Share this project:

Updates

Submission history