-
-
Each claim gets Backed / May conflict / No receipt, with the Wikipedia sentence it relied on
-
When a simpler checker would have answered differently, a note shows why Receipts set that sentence aside
-
In-browser pipeline: split, Wikipedia search, embeddings, NLI, relevance gate, verdict
-
Held-out answers: false "may conflict" on true facts 3.6% vs 12.9% for a simple checker; triage view; limits stated
-
What now? Copy the answer with sources, or a question asking the chatbot to back up the rest
-
Receipts: a receipt for every claim in an AI answer, checked in your browser
In three lines. Receipts checks an AI answer claim by claim against Wikipedia, right in the browser, and always shows the sentence it used. On held-out, human-labelled chatbot answers it wrongly flags a true fact 3.6 % of the time (a simple checker: 12.9 %). It cannot check everything — many claims end as "No receipt" — so it hands you a next step: a draft with sources, or a question to send back to the chatbot.
A student's five minutes (an illustration of the intended use, not a user study). Mina's history homework asks about a singer from the 1920s. She asks a chatbot, gets a confident paragraph, and pastes it into Receipts. Two sentences come back Backed, each with the Wikipedia sentence and a link. One birth year comes back May conflict, next to the Wikipedia sentence that says otherwise. Three come back No receipt. She presses Copy the answer with sources for her notes, then Copy a question for the chatbot to ask for a source for the other three — and checks those herself.
Try it in a minute: open the link → click "Eiffel Tower" (a paragraph with one planted mistake) or a real ChatGPT answer → Check the receipts → under "What now?", copy the answer with sources or a follow-up question for the chatbot. No account, no key; the first visit downloads two small models (about 110 MB, cached afterwards). Measured with the CPU slowed 4× to mimic a school laptop: first result in 21 s, all four claims in 44 s, model download included.
Inspiration (problem statement)
About a quarter of U.S. teens (26 %, double the year before) say they have used ChatGPT for schoolwork, and 54 % say it is fine to use it to research new topics (Pew Research Center, January 2025). Chatbots sound equally sure when they are right and when they are wrong, and checking every sentence by hand is slow, so most people don't. The ML Empowerment curriculum puts it simply: "AI can generate incorrect information… Verify with trusted sources." I wanted a tool that does the slow part of verifying — finding the one sentence that backs or contradicts each claim — and leaves the judgement to the student.
What it does (solution overview + key features)
- Claim by claim: splits the answer into sentences and then into small claims (dates in parentheses, "X was a Y who Z", "…, and she retired …").
- A receipt for every verdict: each claim is Backed, May conflict or No receipt, with the Wikipedia sentence and link. "No receipt" means "check this one yourself", never "false".
- Shows its work: when a simpler checker (top sentence + raw model) would have answered differently, a ⚖︎ note says what it would have said and why Receipts set that sentence aside.
- What now? One click copies the answer with a Wikipedia footnote on every backed sentence and a visible "[check: …]" on the rest — or a follow-up for the chatbot that asks for a checkable source for each claim Receipts could not back (quoting Wikipedia when there is a conflict).
- Free and private: models run in the browser (transformers.js). The answer is never sent as one block and never to an AI company; each claim's text goes to Wikipedia's search, one at a time.
- For teachers: a 15-minute classroom activity built on the app (docs/classroom.md in the repo).
How I built it (technologies)
- Pipeline (all in the browser): split → Wikipedia search per claim (MediaWiki API) → word-overlap pre-filter (40 sentences) →
all-MiniLM-L6-v2embeddings pick the top 5 →nli-deberta-v3-xsmall(8-bit ONNX via transformers.js / WebAssembly, one batched pass) scores backs / contradicts / neutral → decision rules → verdict. - The decision rules are the part I designed. (1) Asymmetric evidence: any good sentence may back a claim, but only the single most relevant one may raise a conflict — on real chatbot answers this is what does the work. (2) Relevance gate: a sentence must be close in meaning, and a conflict must be about the same subject (not "her father…") — this matters when the evidence is about something else.
- No training. Pretrained models only; the only tuned numbers (relevance bar, two thresholds, subject rules) were chosen on development topics with selection rules pushed to GitHub before each run (GitHub Actions timestamps are listed in docs/eval.md), then frozen.
- Engineering: TypeScript, Clean Architecture (pure domain ← use case ← adapters ← UI, checked by a script), 53 unit tests, every per-item result committed, Playwright checks at 390 px and 1280 px, a GitHub Pages deploy that runs tests, typecheck and the layer check first.
How well it works
Full tables, ablation, paired topic-bootstrap intervals and misses: https://github.com/Ryugi62/receipts/blob/main/docs/eval.md
- Held-out chatbot answers (FActScore PerplexityAI biographies, 1,253 human-labelled facts, 35 topics, run once after freezing; new answers and labels, though 30 of the 35 people also appear in earlier test topics): false "may conflict" on facts humans found true 3.6 % vs 12.9 % for the simple checker; balanced accuracy 70.7 % vs 66.6 % (paired difference +4.1 points, 95 % CI 1.9–6.4). The ablation shows the asymmetric rule does the work; without it the gain disappears and false conflicts jump to 23 %.
- As a filter, on those human-split facts: it marks 55.7 % Backed (94.1 % right), and 79.2 % of the facts humans could not support stay in the "check yourself" pile — skipping the same share at random would keep 44.3 %.
- Beyond biographies only (run once with the frozen settings, through the app's whole path — topic guess, splitter, live Wikipedia): 355 FEVER claims on mixed topics (films, bands, places, science, people…) → balanced accuracy 75.5 % vs 68.4 % (+7.0 points, CI 4.6–9.5); false "may conflict" on true claims 5.4 % vs 18.4 %; when it says "May conflict" it is right 88 % of the time.
- Off-topic evidence (easy synthetic test, 712 FEVER pairs): given an unrelated sentence, the small model alone says "contradiction" 74.3 % of the time, 48.2 % with my thresholds but no gate, 0.3 % with the gate.
- The price, plainly: it names few mistakes by itself (6 % of unsupported biography facts, 28 % of false FEVER claims become "May conflict"; on biographies only 24 % of its "May conflict" calls hit a real error). With the correct FEVER evidence handed over, it calls 49 % of refutations a conflict vs 89 % for the simple checker — the cost of refusing to judge on weak evidence. On whole raw ChatGPT paragraphs it ties the simple checker (55.2 % vs 54.9 %, 375 sentences on 49 new topics) and backs only 13 % of the sentences humans support.
Challenges I ran into
- The model said "contradiction" to sentences that weren't about the claim. My first version told me Julia Faye's death date was contradicted by a sentence about her father. Measuring it turned an anecdote into a number and a fix.
- My first fix wasn't good enough. On held-out answers v1 only tied the simplest baseline: any of the top five sentences could raise a conflict. Revision 2 lets only the most relevant sentence do that — and the ablation shows that rule, not my relevance gate, helps on real answers.
- Three later fixes did not work, and I did not ship them. Cleaning the evidence text ("He was…" → the person's name, native-script names dropped, one date format), reading 10 sentences instead of 5, and a richer splitter for whole paragraphs each had to clear a bar on development topics that I pushed to GitHub before running them (+1 point, +1 point, +2 points). They gained 0.5, 0.1 and 0.3. Lesson: on raw paragraphs the gap is retrieval and the model, not my text rules.
- "She" broke retrieval, and long sentences never got "backed" — so pronouns resolve to the topic and long sentences split into smaller claims with tested rules.
Related work, and what is new here
Checking claims against Wikipedia with an inference model is an established research setup (FEVER, FActScore, SummaC/AlignScore-style checkers). Receipts does not claim a new model. What it adds: it runs entirely in a student's browser with no key, account or server; the asymmetric decision rule, measured to cut false "this is wrong" calls on true facts from 12.9 % to 3.6 % on held-out answers; a visible comparison with what a plain checker would have said; and a next step instead of a dead end.
Target users and social impact
High-school and college students who use chatbots for homework, and teachers who want to teach "verify before you trust" with a concrete tool — the classroom activity takes 15 minutes and needs only a browser. It is free, needs no account, keeps the answer away from AI companies, and runs on an ordinary laptop. Teaching point built into the product: a model can be confidently wrong in a measurable way, and the fix is to check what it actually read.
Accomplishments that I'm proud of
- A fix measured on data I hadn't seen, without training anything: false alarms on true facts down from 12.9 % to 3.6 % on held-out answers, and +7 points beyond biographies on the app's whole path.
- Publishing what didn't go my way — v1's tie, three fixes that missed their bar, the tie on raw paragraphs — next to what did.
- A real app, not a notebook: browser-only, no account, 53 unit tests, CI checks, every number on the page generated from committed results.
What I learned
- A model can be right for the wrong reason, and you can measure it — by feeding it sentences that should change nothing.
- Evaluation design is most of the work. Splitting topics before looking, pushing the selection rule before the run, and keeping the held-out run for last changed what I could honestly claim.
- Precision and recall are a product decision. For a student, a false "this is wrong" about a true fact is worse than "check this yourself", so I tuned for trustworthy calls and renamed "Contradicted" to "May conflict" when the numbers said so.
What's next
A better splitter for whole raw paragraphs (the weak spot), other languages (the model has multilingual siblings), and a browser extension that adds receipts next to a chatbot's reply.
What works and what doesn't
Works: English text about people, places, events and other things with a Wikipedia article. Limits: Wikipedia is not the truth, and "No receipt" is not "false"; "May conflict" is a hint; whole raw paragraphs are its weak spot; English only; no user study yet; the biography labels were made on 2023 Wikipedia and the FEVER labels on 2017 Wikipedia, while the app reads today's.
Built during the challenge; tools disclosed
Built solo; all code written during the event (first commit 2026-10-05). An AI coding assistant (Claude) helped write code and text; every number comes from scripts in the repo. Data: Symmetric FEVER (Schuster et al., 2019, CC BY-SA 3.0), FActScore (Min et al., 2023, MIT). Models: nli-deberta-v3-xsmall, all-MiniLM-L6-v2. Narration: Microsoft Edge neural TTS.
Built With
- deberta
- github
- onnx
- playwright
- sentence-transformers
- transformers.js
- typescript
- vite
- vitest
- wikipedia-api
Log in or sign up for Devpost to join the conversation.