Inspiration

AI-generated answers are becoming part of everyday life, but a confident answer is not necessarily a correct one. We wanted to build a system that does more than produce another answer — it should challenge the reasoning behind a claim.

That led us to RealityCheck AI, an evidence-first claim analysis system designed to identify contradictions, verify numerical statements independently, and clearly separate what evidence actually supports from what it does not.

What I built

RealityCheck AI accepts either a claim with supporting evidence or an uploaded PDF document.

For a PDF, the system:

  1. Extracts the document text.
  2. Identifies important factual and testable claims.
  3. Analyzes the claims against the supplied evidence.
  4. Independently validates numerical assertions using Python.
  5. Identifies contradictions, unsupported claims, and evidence boundaries.
  6. Produces a claim-by-claim document report and summary.

For example, when a claim says:

"Students who use AI score 40% higher in exams."

and the supplied evidence reports average scores of 78 versus 65, RealityCheck calculates the increase independently:

(78 - 65) / 65 × 100 = 20%

The system therefore flags the 40% claim as contradicted by the supplied evidence.

How it works

RealityCheck combines language-model reasoning with deterministic validation.

The language model is used for tasks such as extracting claims and identifying explicit and unsupported statements. Numerical comparisons are independently calculated using Python rather than trusting the model's arithmetic.

This separation is important: the model interprets the evidence, while deterministic code verifies numerical relationships.

Challenges

One of the biggest challenges was dealing with unreliable model output formats. Language models can return values such as "40" instead of the numeric value 40. We added validation and type conversion so numerical verification remains deterministic.

We also had to establish an evidence boundary so the system would not automatically turn a correlation or observation into a causal conclusion.

What we learned

We learned that trustworthy AI systems should not rely on a single model response. Combining language understanding with independent checks makes the reasoning easier to inspect and the result easier to challenge.

Future scope

Future versions could support more document formats, richer citation and source tracking, stronger claim classification, cross-document comparison, and deeper research-level verification.

Built With

Share this project:

Updates

Submission history