## Track: Access to Justice & Civic Tech
Inspiration
I am a student in Nigeria. Everyone I know has a story: a police officer demanding to search a phone, a landlord giving seven days to leave, a salary that stops coming, a loan app threatening to message every contact. Almost nobody involved has ever spoken to a lawyer. The laws that protect people exist, but they live in scanned PDFs written for lawyers.
Chatbots do not fix this. Ask a general model about Nigerian tenancy law and it will confidently cite a section that does not exist. In law, a wrong answer delivered confidently is worse than no answer. So the question I set out to answer was: can an AI legal guide be built so that it cannot invent the law?
What it does
Wetin takes a question in English or Nigerian Pidgin and answers with what the written law actually says.
- Understands the question and rewrites it into statutory search terms. Pidgin like "dem no gree give bail" becomes "bail at police station" and "personal liberty".
- Retrieves the relevant sections from a corpus of 1,460 sections across eight statutes: the 1999 Constitution, the Police Act 2020, the Administration of Criminal Justice Act 2015, the Nigeria Data Protection Act 2023, the Labour Act, and the Lagos Tenancy Law 2011, the Cyber crimes Act 2015, the Child's Rights Act 2003.
- Answers only from those sections. Every claim carries a citation chip that jumps to the section text. Citations to anything outside the retrieved set are dropped on the server before the answer reaches the user.
- Verifies every quote. Each quoted excerpt is matched against the real statute text and highlighted inside the full section. If the model paraphrased, the quote is snapped to the exact statutory wording. If no match exists, it is flagged.
- Gives next steps and drafts the document the person actually needs: a letter to a landlord, a petition to the Police Service Commission, a complaint to the Nigeria Data Protection Commission, a letter to an employer.
- Flags urgency. If someone is in custody right now, it surfaces the constitutional 24 and 48 hour rule and links to free legal help.
How I built it
- Corpus: I downloaded the gazette PDFs, extracted the text, and wrote a Python parser that finds each statute's body, walks the sections in order so schedules and tables do not confuse it, pulls titles from the arrangement of sections, and fixes common OCR errors. The output is one JSON file of sections with source links.
- Retrieval: BM25 search with MiniSearch, running in process and chunked by subsection. A fast model then reranks the top 24 candidates. There is no vector database, so the whole app is a single Next.js deploy.
- Generation: three model calls through Groq. Qwen 27B handles query expansion and reranking. GPT-OSS 120B writes the answer as strict JSON with citation tokens and verbatim quotes.
- Grounding layer: server-side code filters citations to the retrieved set, verifies quotes against the section text, and snaps near matches to the exact wording. The interface shows the verification count on every answer.
- Resilience: a fallback chain across three models handles rate limits, timeouts and network errors, so the free tier stays usable under load.
- Frontend: Next.js 16, React 19 and Tailwind 4. Mobile first, light and dark themes, English and Pidgin, with streamed status so the person sees the sections being read before the answer arrives.
Frontend: Next.js 16, React 19 and Tailwind 4. Mobile first, light and dark themes, English and Pidgin, with streamed status so the person sees the sections being read before the answer arrives.
Tools and credits: built with Claude Code as an AI coding assistant. Narration in the demo video uses a synthetic voice. Statute texts come from public PDFs published by PLAC, SabiLaw, the Policing Law database and DataGuidance.
Challenges I ran into
The law is locked in scans. Three of the six statutes were image-only PDFs. Finding text copies and cleaning OCR artefacts took a large part of the build.
Section parsing is not uniform. Each gazette formats sections differently: numbers on their own line, marginal notes mixed into the text, three-column tables, schedules that restart numbering at 1. One statute had to be parsed from a layout-preserving extraction because the default scrambled its columns.
Pidgin retrieval. Keyword search on "dem carry my brother go station" returns nothing useful. Query expansion by a small model fixed this and made English and Pidgin equally usable.
Keeping the model honest. Open models paraphrase quotes even when told not to. Prompting alone was not enough. The fix was code that checks every quote and replaces paraphrases with the exact statute text.
Free-tier limits. Groq allows 8,000 tokens per minute per model. I had to fit the full answer prompt inside that budget and build fallback across models.
Accomplishments that I'm proud of
- In the recorded live demo, every quote was verified against the statute text: 4 of 4 in English and 3 of 3 in Pidgin.
- The same question works in English and in Pidgin, and the answer comes back in the language chosen.
- A working letter generator that cites only sections the person has just read.
- The whole pipeline runs on one Next.js deploy, with no vector database, on a free model tier.
What I learned
Grounding is a systems problem, not a prompting problem. The prompt asks for verbatim quotes, but what makes the product trustworthy is the plain code that checks them. I also learned that people do not ask legal questions in legal language. A cheap translation step before retrieval matters more than a bigger model after it.
What's next for Wetin
- More statutes: the Criminal Code, Penal Code, state criminal justice laws, Cybercrimes Act, Child Rights Act, and tenancy laws of other states.
- Answers in Hausa, Yoruba and Igbo.
- A WhatsApp interface, since that is where Nigerians already are.
- An offline rights card for police stops that works with no data.
- Partnerships with the Legal Aid Council and campus legal clinics to route people to real help.
Built With
- gpt-oss
- groq
- minisearch
- next.js
- pdftotext
- python
- qwen
- react
- tailwindcss
- typescript
- vercel
Log in or sign up for Devpost to join the conversation.