Inspiration

You sign up for internet at ₹499 a month with free installation.

Two months later, an ₹849 bill arrives with a ₹350 installation fee.

You know it is wrong. But the proof is buried in an old email, the support process will take hours, and even if the provider approves a credit, you still have to remember to check whether it ever arrived.

Most people eventually decide that ₹350 is not worth the chase. The same thing happens with refunds, warranties, bookings, deliveries, and subscription promises every day.

Companies have systems that reconcile what they expected with what actually happened. People do not.

RealityCheck is the agent I wanted on my side: one that remembers the promise, notices when reality changes, and stays on the case until the outcome is proven.

What it does

RealityCheck handles one complete job from start to finish:

  1. Remember the promise. It turns emails, receipts, bills, and screenshots into clear, evidence-backed expectations.
  2. Watch what happens next. When a new bill or outcome arrives, it compares reality with the original agreement.
  3. Explain the difference. It shows the exact mismatch and the source evidence behind it.
  4. Ask before acting. It cannot contact a provider without explicit, case-scoped approval.
  5. Take the next step. Once approved, it prepares and sends an evidence-backed correction request.
  6. Keep following up. A provider saying “approved” creates a new promise for RealityCheck to monitor.
  7. Close only with proof. The case ends only when fresh evidence confirms the correction actually happened.

It is not a chatbot that tells you what to do. It owns the workflow, waits between events, takes permissioned action, and verifies the result.

The demo story

The demo uses FiberMax, a clearly labeled fictional provider sandbox.

FiberMax promises ₹499/month with free installation. Later, it sends an ₹849 bill containing a ₹350 installation charge.

RealityCheck:

  • finds the exact ₹350 mismatch;
  • cites the sentence promising free installation;
  • blocks provider contact until the user gives scoped approval;
  • sends the approved correction;
  • records FiberMax’s promise to issue a credit within 48 hours;
  • keeps the case open while that credit is still owed; and
  • verifies the adjustment evidence before reporting ₹350 recovered.

That last step matters. “Your credit was approved” is not the same as getting the credit. RealityCheck understands the difference.

Why this is agentic

RealityCheck owns a long-running goal: make reality match the agreement.

A Google ADK specialist fleet divides the work across Expectation, Observation, Judge, Guardian, Resolution, OWED, and Outcome agents. The system decides which specialist should act, preserves state between events, pauses for human approval when required, and resumes when new evidence arrives.

The user does not have to repeatedly explain the case or remember the next follow-up.

How I built it

  • Gemini 3.5 Flash reads messy human agreements and extracts structured terms with exact supporting quotes.
  • The Google Gen AI SDK provides typed structured output.
  • Google ADK defines and coordinates the specialist agent fleet.
  • Deterministic Python owns arithmetic, exact comparisons, state transitions, and permission checks so the model cannot invent numeric truth.
  • Google Cloud Firestore stores cases, evidence references, audit events, and unresolved obligations across sessions.
  • FastAPI runs the state machine and API.
  • The public dashboard runs on Vercel, while durable production state lives in Google Cloud Firestore.

Trust is part of the product

A consumer agent should not become another thing the user has to fear.

RealityCheck therefore keeps three boundaries visible:

  • Evidence before claims: every important conclusion points back to its source.
  • Permission before action: the independent Guardian Agent enforces scoped approval.
  • Proof before closure: a promise of correction never counts as a completed outcome.

Evidence is hashed, audit events are chained, uncertainty remains visible, and repeated events are handled safely instead of duplicating actions.

What is real

The public demo is a working product, not a collection of mock screens.

  • Gemini 3.5 Flash is connected in production.
  • Case state persists in Google Cloud Firestore across sessions.
  • The full capture → compare → approve → act → monitor → verify lifecycle runs through the live dashboard.
  • The /api/health endpoint exposes the actual model, Firestore project, database location, and sandbox boundary.
  • FiberMax is fictional so the public demo never contacts or misrepresents a real company.

Public app: https://realitycheck-agent.vercel.app

Repository: https://github.com/vivekyarra/RealityCheck

Challenges I ran into

The subtraction was easy. Trustworthy judgment was hard.

A higher bill is not automatically wrong: taxes, prorations, upgrades, partial deliveries, or matching credits can explain a difference. I had to separate semantic understanding from deterministic truth. Gemini interprets unstructured language; code owns the math and permission boundaries.

I also learned that resolution is a timeline, not a button. The system needed durable memory so an approved credit remained an open obligation until later evidence proved it arrived.

What I am proud of

I built the complete loop instead of stopping at anomaly detection.

RealityCheck can move from a buried promise to a verified outcome while showing the user exactly what it knows, what it is waiting for, and what it is allowed to do. The same design can extend beyond billing to refunds, warranties, travel promises, specifications, deliveries, and deadlines.

What I learned

The best agent is not the one that talks the most. It is the one that quietly owns an annoying job, acts safely, and comes back with proof.

Models are excellent at understanding messy human language. Deterministic systems are better at arithmetic, permissions, and closure rules. RealityCheck became more reliable when I let each do the work it is best suited for.

What's next

Next I want to add consented connectors for email, receipts, calendars, and provider APIs; reusable policy packs for common consumer promises; scheduled background observation; and a privacy-preserving personal evidence vault.

The long-term goal is simple:

When reality drifts from what you were promised, you should not have to notice, calculate, chase, and remember alone.

Built With

  • fastapi
  • gemini-3-5-flash
  • google-adk
  • google-cloud-firestore
  • google-genai-sdk
  • python
  • vercel
Share this project:

Updates

Submission history