Inspiration
I don't have a film background. I just like movies — which means I'm on wikis constantly, checking who was in what, when something came out, how one film connects to another.
Enough times I found something that was simply wrong. Not vandalism, and never obvious: a fact that used to be true and quietly stopped being true, sitting there looking exactly like every correct sentence around it. Nobody had edited the page. The world had just moved on.
That's where Continuity came from — an agent that reads a page fact by fact and re-checks each one against the live web, using Parallel to do the research.
The review part came from my background in code. I read git diffs all day, and it's still the clearest way I know to say here's what it said, here's what it should say. So the agent never decides anything. It puts both versions side by side with its sources, and the maintainer makes the call.
What it does
You open a wiki article and press one button. That's the whole trigger — no setup, no config, no schedule to configure.
From there the agent works through the page:
- Reads it and writes down every factual claim it makes, one fact at a time.
- Decides what's worth checking right now. Claims it verified last week get skipped.
- Researches each one on the live web through Parallel.
- Judges what it found with Gemini — still true, changed, or the sources disagree.
- Writes the edit for anything that moved, with the citation attached.
- Checks that edit against the rest of the page, so fixing one paragraph doesn't contradict another.
- Hands it back as a diff — old text, new text, sources, confidence score.
Then it stops and waits for you. Approve it, rewrite it, or throw it out. Nothing reaches the page until a person says so.
How we built it
Python on the backend, plain JavaScript on the front — no framework, no build step. Gemini does the judging, Parallel does the research, and it all runs as one Cloud Run service that scales to zero.
The part I spent most of the time on isn't the model calls. It's the ledger: a small Python core that tracks every claim, when it was last checked, when it's due next, and what else it's connected to. It has no dependencies and never touches the network, so all the interesting logic is testable without a cloud account or an API key. The vendor SDKs live only at the edges.
Challenges we ran into
Every run costs money. One pass over a single page is a web search per claim plus several Gemini calls. That's fine once. It is not fine when you're tuning a prompt and running it forty times in an afternoon.
So I recorded fixtures. The first live run wrote every search result and model response to disk, and after that the whole pipeline could replay from them for free. That's what made the thing developable — I could rewrite the classify stage, run all six stages end to end, and spend nothing.
I built the rest locally for the same reason. The ledger core has no dependencies at all, so most of the test suite runs on a bare Python interpreter with nothing installed — no database, no emulator, no cloud project. The billing meter didn't start until the day I actually deployed.
Accomplishments that we're proud of
The whole thing cost about $10 to build. Recording fixtures early meant nearly all of the development replayed for free — every prompt rewrite, every end-to-end run, the whole test suite. The credits went where they actually mattered: hosting, Firestore, and real Gemini and Parallel calls for the demo, instead of being burned on debug run of a stage I was still rewriting. Getting that discipline in place on day one is the reason there was budget left to deploy at all.
What we learned
Plan further ahead than feels necessary. I rewrote a lot of this project, and almost none of it was because I learned something I couldn't have known upfront — it was because I started building before the idea was settled.
Log in or sign up for Devpost to join the conversation.