-
-
A reasoning-first interface for checking claims against evidence.
-
Upload a document, extract claims, validate them, and see a claim-by-claim evidence report.
-
Deterministic Python validation catches a 40% claim that the evidence supports at only 20%.
-
Responsive interface designed for desktop and mobile workflows.
Inspiration
AI-generated answers are becoming part of everyday life, but a confident answer is not necessarily a correct one. We wanted to build a system that does more than produce another answer — it should challenge the reasoning behind a claim.
That led us to RealityCheck AI, an evidence-first claim analysis system designed to identify contradictions, verify numerical statements independently, and clearly separate what evidence actually supports from what it does not.
What I built
RealityCheck AI accepts either a claim with supporting evidence or an uploaded PDF document.
For a PDF, the system:
- Extracts the document text.
- Identifies important factual and testable claims.
- Analyzes the claims against the supplied evidence.
- Independently validates numerical assertions using Python.
- Identifies contradictions, unsupported claims, and evidence boundaries.
- Produces a claim-by-claim document report and summary.
For example, when a claim says:
"Students who use AI score 40% higher in exams."
and the supplied evidence reports average scores of 78 versus 65, RealityCheck calculates the increase independently:
(78 - 65) / 65 × 100 = 20%
The system therefore flags the 40% claim as contradicted by the supplied evidence.
How it works
RealityCheck combines language-model reasoning with deterministic validation.
The language model is used for tasks such as extracting claims and identifying explicit and unsupported statements. Numerical comparisons are independently calculated using Python rather than trusting the model's arithmetic.
This separation is important: the model interprets the evidence, while deterministic code verifies numerical relationships.
Challenges
One of the biggest challenges was dealing with unreliable model output formats. Language models can return values such as "40" instead of the numeric value 40. We added validation and type conversion so numerical verification remains deterministic.
We also had to establish an evidence boundary so the system would not automatically turn a correlation or observation into a causal conclusion.
What we learned
We learned that trustworthy AI systems should not rely on a single model response. Combining language understanding with independent checks makes the reasoning easier to inspect and the result easier to challenge.
Future scope
Future versions could support more document formats, richer citation and source tracking, stronger claim classification, cross-document comparison, and deeper research-level verification.
Log in or sign up for Devpost to join the conversation.