Inspiration

Nigeria's tax landscape changed fundamentally with the Nigeria Tax Act 2025 (effective 2026): new bands, new rent-relief rules, new obligations. But accurate guidance didn't get any closer — accountants are expensive, official tools are scarce, and internet access is unreliable or unavailable in much of the country. I kept meeting people who wanted to know simply: "How much tax will I actually pay?"

That question shouldn't require connectivity, a subscription, or a lawyer. It should be answerable offline, on an ordinary laptop, in the language people actually speak.

What it does

TaxSabi is a fully offline Nigerian personal-income-tax assistant. Ask in English or Nigerian Pidgin and it answers with band-by-band calculations, statutory citations, relief handling (rent, pension, NHF, NHIS, mortgage interest, life insurance), monthly-to-annual conversion, and what-if scenarios. It defaults to Nigeria/2026 when you omit them and declines other jurisdictions and years instead of guessing.

It ships as a single 1.70 GiB GGUF (Q8_0) file that runs through llama.cpp — no internet, no GPU, no install beyond downloading the model.

How we built it

  • A deterministic rules engine came first. A Python Decimal tax engine implements the Nigeria Tax Act 2025 (bands, reliefs, evidence rules), with every legal fact traced in a source register (F-001–F-019).
  • The engine wrote the textbook. Every calculation in the training data was engine-computed: a 695K-character domain corpus for continued pretraining, 1,929 engine-verified SFT examples, 3,144 binary-labelled KTO examples (engine/citation rules), and 1,284 engine-grounded GRPO prompts.
  • Four-stage fine-tuning of Qwen3-1.7B with Unsloth QLoRA on an A100: DAPT → SFT → KTO → GRPO, no thinking mode. SFT taught working-first answers (total last); GRPO used the rules engine as the reward (exact final tax = 1.0, exact chargeable income = 0.25).
  • We treated quantization as a correctness gate. We exported Q4_K_M, Q5_K_M, Q6_K and Q8_0 and tested each through llama.cpp against the full-precision model: Q4/Q6 corrupted the 2026 band table (and Q4 flipped arithmetic), Q5 substituted an older table. Q8_0 reproduced the full-precision model, so we shipped it and accepted the throughput cost.
  • We evaluated adversarially: held-out novel-amount probe, dev probe, a 100-prompt paraphrase suite, and Pidgin cases — all engine-scored.

Challenges we ran into

  • Our first SFT format put the headline total first; the model then produced totals that contradicted its own band breakdown. We reversed the format (working first, total last) and re-measured.
  • A targeted SFT top-up and a third GRPO iteration both regressed behaviour without improving accuracy — our pre-agreed gates caught both, and we froze the better checkpoint (SFT v2, GRPO v2) instead of shipping a "newer" model.
  • Quantization: Q4_K_M measured noticeably worse than the full-precision model on statutory knowledge. Documenting and acting on that cost us throughput, but the shipped artifact is the GGUF — accuracy had to win.
  • Band-list answers are phrasing-sensitive; 2 of 6 tested phrasings fail even at full precision. We publish this failure mode rather than hide it.

Accomplishments that we're proud of

-- Every training number is engine-guaranteed — no AI-generated arithmetic anywhere in the dataset, and a verifier rejects any record with unsourced amounts.

  • Shipped Q8_0 results: held-out novel calculations 5/10 exact (7/10 chargeable income), dev probe 8/9, paraphrase suite 17/40 exact with 5/8 phrasing groups consistent, and zero out-of-register citations across 1,010 captured outputs.
  • Participant-mode profiler on the shipped GGUF: 3.3–4.7 t/s generation, ~1.94 GB peak RSS, no throttling — far inside the 7 GB budget, fully offline (dev laptop, slower than the Standard Laptop spec).
  • Full Gate-2 provenance: the final adapter (Git LFS), training scripts, per-stage logs and per-step GRPO metrics, dataset documentation, SHA256 checksums and the merge/quantization study — all in provenance/.

What's next for TaxSabi

  • Fact-ledger architecture (post-competition): grammar-constrained extraction plus the rules engine so multi-turn conversations accumulate facts safely — language from the LLM, arithmetic from code that cannot hallucinate.
  • Desktop bundles rebuilt for the final model (Windows/Linux/macOS), then a one-click installer.
  • Mobile — the model runs in under 2 GB of RAM; the work is packaging llama.cpp for Android/iOS.
  • More Nigerian languages — Yoruba and Hausa, with the same human-review standard held for Pidgin.

Built With

  • adtc-profiler
  • bash
  • decimal
  • gguf
  • google-colab
  • html
  • huggingface
  • javascript
  • llama-bench
  • llama.cpp
  • nigeria-tax-act-2025
  • peft
  • python
  • qlora
  • qwen2.5
  • unsloth
Share this project:

Updates

Submission history