Inspiration
Nigeria's tax landscape changed fundamentally with the Nigeria Tax Act 2025 (effective 2026): new bands, new rent-relief rules, new obligations. But accurate guidance didn't get any closer — accountants are expensive, official tools are scarce, and internet access is unreliable or unavailable in much of the country. I kept meeting people who wanted to know simply: "How much tax will I actually pay?"
That question shouldn't require connectivity, a subscription, or a lawyer. It should be answerable offline, on an ordinary laptop, in the language people actually speak.
What it does
TaxSabi is a fully offline Nigerian personal-income-tax assistant. Ask in English or Nigerian Pidgin and it answers with band-by-band calculations, statutory citations, relief handling (rent, pension, NHF, NHIS, mortgage interest, life insurance), monthly-to-annual conversion, and what-if scenarios. It defaults to Nigeria/2026 when you omit them and declines other jurisdictions and years instead of guessing.
It ships as a single 1.70 GiB GGUF (Q8_0) file that runs through llama.cpp — no internet, no GPU, no install beyond downloading the model.
How we built it
- A deterministic rules engine came first. A Python
Decimaltax engine implements the Nigeria Tax Act 2025 (bands, reliefs, evidence rules), with every legal fact traced in a source register (F-001–F-019). - The engine wrote the textbook. Every calculation in the training data was engine-computed: a 695K-character domain corpus for continued pretraining, 1,929 engine-verified SFT examples, 3,144 binary-labelled KTO examples (engine/citation rules), and 1,284 engine-grounded GRPO prompts.
- Four-stage fine-tuning of Qwen3-1.7B with Unsloth QLoRA on an A100: DAPT → SFT → KTO → GRPO, no thinking mode. SFT taught working-first answers (total last); GRPO used the rules engine as the reward (exact final tax = 1.0, exact chargeable income = 0.25).
- We treated quantization as a correctness gate. We exported Q4_K_M, Q5_K_M, Q6_K and Q8_0 and tested each through llama.cpp against the full-precision model: Q4/Q6 corrupted the 2026 band table (and Q4 flipped arithmetic), Q5 substituted an older table. Q8_0 reproduced the full-precision model, so we shipped it and accepted the throughput cost.
- We evaluated adversarially: held-out novel-amount probe, dev probe, a 100-prompt paraphrase suite, and Pidgin cases — all engine-scored.
Challenges we ran into
- Our first SFT format put the headline total first; the model then produced totals that contradicted its own band breakdown. We reversed the format (working first, total last) and re-measured.
- A targeted SFT top-up and a third GRPO iteration both regressed behaviour without improving accuracy — our pre-agreed gates caught both, and we froze the better checkpoint (SFT v2, GRPO v2) instead of shipping a "newer" model.
- Quantization: Q4_K_M measured noticeably worse than the full-precision model on statutory knowledge. Documenting and acting on that cost us throughput, but the shipped artifact is the GGUF — accuracy had to win.
- Band-list answers are phrasing-sensitive; 2 of 6 tested phrasings fail even at full precision. We publish this failure mode rather than hide it.
Accomplishments that we're proud of
-- Every training number is engine-guaranteed — no AI-generated arithmetic anywhere in the dataset, and a verifier rejects any record with unsourced amounts.
- Shipped Q8_0 results: held-out novel calculations 5/10 exact (7/10 chargeable income), dev probe 8/9, paraphrase suite 17/40 exact with 5/8 phrasing groups consistent, and zero out-of-register citations across 1,010 captured outputs.
- Participant-mode profiler on the shipped GGUF: 3.3–4.7 t/s generation, ~1.94 GB peak RSS, no throttling — far inside the 7 GB budget, fully offline (dev laptop, slower than the Standard Laptop spec).
- Full Gate-2 provenance: the final adapter (Git LFS), training scripts, per-stage logs and per-step GRPO metrics, dataset documentation, SHA256 checksums and the merge/quantization study — all in
provenance/.
What's next for TaxSabi
- Fact-ledger architecture (post-competition): grammar-constrained extraction plus the rules engine so multi-turn conversations accumulate facts safely — language from the LLM, arithmetic from code that cannot hallucinate.
- Desktop bundles rebuilt for the final model (Windows/Linux/macOS), then a one-click installer.
- Mobile — the model runs in under 2 GB of RAM; the work is packaging llama.cpp for Android/iOS.
- More Nigerian languages — Yoruba and Hausa, with the same human-review standard held for Pidgin.
Built With
- adtc-profiler
- bash
- decimal
- gguf
- google-colab
- html
- huggingface
- javascript
- llama-bench
- llama.cpp
- nigeria-tax-act-2025
- peft
- python
- qlora
- qwen2.5
- unsloth
Log in or sign up for Devpost to join the conversation.