About the project

Inspiration

Every "AI demo video" tool on the market records something and narrates over it. None of them can answer the only question that matters six weeks later: which specific claim in this finished video is now false?

We watched this fail in a boring, expensive way — a sales deck quoting a price that had changed two releases ago, still being sent to prospects, because nobody re-watches a recording to audit it. The video is not the liability. The unverified claim inside it is.

So we stopped trying to make a better video generator and built a system where an AI is structurally not allowed to say anything it cannot point at.

What it does

Give Veridemo a URL. It opens your product in a real browser, walks the flows a user would take, and produces a narrated demo video in which every spoken claim traces back to a captured page element. Any claim it cannot verify is dropped before the video ships.

The mechanism is a two-part hash:

  • A fact id = hash(url + css_selector) — a stable pointer to one specific claim.
  • A content hash = hash(what that element says right now).

Drift = same fact id, different content hash.

That split is the whole product. Re-crawl the product later and Veridemo knows exactly which claims went false, which video scenes cited them, and which need rebuilding. One changed price rebuilds one scene — not the whole video.

How we built it

Six agents run in sequence behind a FastAPI orchestrator with priority queues and persistent job state:

# Agent Does
1 Explorer Drives real Chromium via Playwright; clicks, types, screenshots, and hashes every fact on the page
2 Knowledge Builds the truth set and a navigation graph from the crawl's actual edges
3 Docs Writes only claims it can cite
4 QA Fails the build on an unsupported number or a dangling citation
5 Demo Renders the narrated video from verified scenes only
6 Release Re-crawls later, diffs the hashes, rebuilds only what drifted

Where the AI sits, and where it deliberately does not. A language model writes the narration and chooses which facts each scene asserts. A neural voice speaks it. But the model is never what decides whether a claim is true: everything it produces is re-checked against the hashed truth set before it can reach a scene, and a mismatch fails verification instead of shipping. The model proposes; the hash decides.

With no API key at all, the pipeline still runs end to end on a deterministic matcher — so a judge can clone the repo and reproduce every number in our submission without signing up for anything.

harness/veridemo_mcp.py also exposes the pipeline as MCP tools (crawl_product, list_facts, check_narration, review_script, publish_demo), so an LLM agent composes the pipeline itself rather than running a fixed script. publish_demo is annotated destructive and held for human approval.

Challenges we ran into

The guard has to be able to actually fail. Our first QA agent passed everything, which made it decorative. We rewrote it to check cited numbers against the source: a hallucinated $129 against a page that says $99 now produces status: FAILED. An adversarial audit of our own guard found and fixed two real holes (a broken numeric check and a broken transport).

LLM segmentation is not free. Doing claim-segmentation and fact-matching as two separate passes throws away the context that makes the match good. Combining them into one pass improved matching sharply — but then an LLM mismatch could ship, so every extracted claim still runs through the same numeric-support guard afterwards.

Reproducibility. A demo video is worthless as evidence if the judges can't tell whether it was staged. So our submission video is generated by a committed script that reads the same committed artifacts as the pipeline — not edited by hand.

Accomplishments that we're proud of

A real, unattended run against https://httpbin.org with no prior data:

  • 8 pages crawled, 356 real DOM elements hashed, 8 real click actions executed
  • A 3-scene narrated MP4 rendered with AI voiceover
  • QA passed 11/11 checks (100%) — every claim verified against its citations, zero dangling citations, zero uncited claims

All of it ships in the repo as evidence (docs/proof-run/), including the rendered video, the full pass/fail report, and the raw screenshots.

And the submission video itself is honest by construction: every frame is either a generated slide or a real artifact, and the actual product-rendered MP4 plays inside it, unedited. Run python tools/build_submission_video.py and you get our video back.

What we learned

Verification isn't a feature you bolt onto a generator — it changes what the unit of work is. Once a sentence is bound to a fact id, "is this still true?" stops being a judgement call and becomes a hash comparison. That one shift is what makes selective rebuilds, drift alerts, and auditable demos all fall out of the same mechanism.

We also learned to distrust our own pass rates. A guard that never fails is not a guard.

What's next for Veridemo

  • Re-crawl on deploy (CI hook), so drift is caught at release time rather than discovered
  • Per-segment demo variants generated from one truth set
  • Broadening the guard past numeric support to entity and capability claims
  • Deeper integration with our sibling submission, Veridemo Watch, which applies the same engine to content a creator has already published

Sibling project disclosure

We are also submitting Veridemo Watch to this hackathon — same team, same verified-fact engine, different problem. Veridemo generates verified demos; Watch audits content that already exists. Different input, different user, shared mechanism, disclosed plainly rather than presented as unrelated.

Video demo link

Upload docs/submission/veridemo_submission.mp4 (3:51) to YouTube as unlisted and paste the URL here https://youtu.be/LCSzXsl73Zs.

Built With

Share this project:

Updates

Submission history