Inspiration

Before a film can be insured, somebody reads the screenplay by hand and checks every invented name against the real world. Every character. Every business. Every in-world title. If a name collides with something real, the production gets sued, or the Errors & Omissions underwriter simply refuses to insure the film.

That report costs $1,399 to $5,000 per feature and takes one to two weeks. It is a human, reading a script, looking for collisions.

When it goes wrong it is not a rounding error. Tarnation was shot for $218 and ran up a $400,000 clearance bill. The Ella Fitzgerald documentary spent 60% of its budget on clearance.

We also looked hard at what everyone else builds for a "cinema + AI" prompt. Every recurring idea is creative generation: storyboard generators, script-to-video pipelines, auto-trailers, AI dubbing. One hackathon had four separate teams building the same highlights generator. Almost nothing touches production operations, meaning scheduling, compliance, rights and clearance, which is exactly where the expensive, unglamorous friction lives.

What it does

Paste a scene. Clearance Desk returns a clearance report: every clearable name, what real thing it collides with, a risk tier, the recommended action, and the SQL query that produced each finding so you can re-run it and check the claim yourself.

The core idea is that clearance asks two questions, and most name-matching answers only the first:

  1. Existence: does something real already carry this name?
  2. Exposure: is it prominent enough to notice, and to sue?
Question Source Scale
Existence, real people imdb.actors ⋈ imdb.roles 817,718 people
Existence, prior titles imdb.movies 388,269 films
Existence, organisations youtube.youtube ~537M rows scanned
Exposure, current attention wiki.wikistat 638,111,114,033 rows, updated daily

Separating them is what makes the risk score reflect now rather than the age of a snapshot.

It also produces the non-obvious results. "Jack Dawson" has almost no screen credits but 32,061 Wikipedia pageviews across 12 languages, because it is the lead character in Titanic. A famous fictional name carries clearance risk too, and a person-database lookup misses it entirely.

Exposure is not one number either. The same table answers two questions no name-lookup can:

  • Where the attention is. JACK DAWSON is 46.6% Spanish, 37.4% Portuguese, only 14.6% English, which makes it a bigger problem for a Latin American release than an Anglophone one.
  • Which way it is moving. JURASSIC PARK attention is +40.2% over 90 days against the prior 90. A name growing more famous is a name getting riskier to use.

How we built it

A deterministic, multi-step agent built on the Google Agent Development Kit.

Google ADK SequentialAgent "clearance_pipeline"
  ├── LlmAgent "element_extractor"   Vertex AI gemini-2.5-flash, temperature 0,
  │                                  typed output_schema → clearable elements
  └── LlmAgent "clearance_analyst"   FunctionTool → official mcp-clickhouse server

then deterministic Python:
  SEARCH  collision queries  ─┐
  EXPOSE  territory + trend  ─┼─→ official mcp-clickhouse MCP server → ClickHouse
  SCORE   derived thresholds ─┘
  REPORT  findings + the SQL for each one

ADK's workflow agents are the right primitive for a brief asking for a deterministic, multi-step agent: with a SequentialAgent the control flow is fixed by the agent graph rather than chosen by a model at runtime, so only the reasoning inside each step is model-driven. That is what makes the same screenplay produce the same report.

Two design decisions followed 2026 evidence rather than instinct:

  • No multi-agent debate. At equal token budgets, single-agent systems match or beat multi-agent on multi-hop reasoning, and debate agents converge to consensus rather than deliberating.
  • No self-critique step. Ungrounded self-critique measurably hurts: one study saw accuracy fall from 98% to 57%. Every check here is grounded in rows returned by ClickHouse, never in the model's opinion of its own output.

The risk thresholds are derived, not chosen: a probe set of known-high names (Taylor Swift 10.6M pageviews/yr, Coca-Cola 2.4M), known-mid (Wetherspoons 180k, Jack Ryan 73k) and known-low (Blue Sun 923, invented names 0) was measured, and the boundaries placed where those groups actually separate.

The masthead texture and logo were generated with Gemini 2.5 Flash Image on Vertex AI and ship as static assets.

Challenges we ran into

The database returned partial answers as complete. The public ClickHouse demo cluster runs max_rows_to_read = 1e9 with read_overflow_mode = break, which stops reading and returns what it has, with no error. A count over the 638-billion-row table came back 638× too low, and the date range was wrong too. Nothing raised. Every statement now forces read_overflow_mode='throw'.

.title() mangles acronyms, and Wikipedia's redirects hid it. Screenplays use capitals, so names were title-cased to reach Wikipedia, but IBM became Ibm. Because Wikipedia redirects Ibm → IBM, the lookup did not return zero; it returned a plausible small number:

Name Naive lookup Found Real Under-reported
IBM Ibm 20,508 1,618,996 79×
BBC Bbc 39,313 2,881,602 73×
McDonald's Mcdonald's 7,637 2,323,669 304×
RONALD MCDONALD Ronald_Mcdonald 228 368,410 1,616×

All scored MEDIUM. All are CRITICAL. The tool was systematically under-warning about exactly the brands most likely to sue.

A shared hourly quota that would have killed the demo during judging. 60 queries/hour per IP, and only 20 of the same query shape. A scan cost 14 queries, so roughly four runs exhausted the budget. Batched to 8, cached to 0 on repeat, with a warm cache shipped in the image.

A legacy storage convention. imdb.movies files titles with the article at the end, so The Godfather is filed as Godfather, The. Querying only the natural form silently missed exact matches; "The Gathering" went from 0 exact matches to 3.

/healthz is reserved by the Google Front End and never reaches the container. The identical container served it correctly on localhost.

ADK needed two things pinned down: it builds its own genai client and defaults to the Gemini Developer API until GOOGLE_GENAI_USE_VERTEXAI=TRUE selects Vertex; and it rejects response_schema inside generate_content_config, because the schema belongs on LlmAgent(output_schema=...) as a pydantic model.

Accomplishments that we're proud of

  • We found thirteen bugs in our own work, and fixed all of them. Twelve of them were the quiet kind that nobody notices until it matters.

  • The tool catches a name we would never have thought to check. "Jack Dawson" has almost no acting credits, so a normal database search shrugs at it. But it is the lead character in Titanic, and 32,061 people looked it up on Wikipedia last year. That is a real problem for a real production, and finding it is the moment the whole idea justified itself.

  • We built it so that when one half breaks, the other half still catches the problem. Our film title list turned out to be incomplete: Jurassic Park is missing from it entirely. The tool still flags Jurassic Park as high risk, because it measures public attention separately. We designed the two checks not to depend on each other, and that decision quietly saved us.

  • You can check our homework. Every single finding shows the exact query behind it. A clearance report is only worth what you can verify, so we made verification the default rather than something you have to ask for.

  • We put the tool's own weaknesses on screen. Where it genuinely cannot tell whether a name refers to a person or a place, it says so and tells you to confirm. It would have been easy to hide that. Admitting it makes the rest of the report worth trusting.

What we learned

  • Being wrong loudly is a gift. Being wrong quietly is what hurts people. A tool that crashes gets fixed in an hour. A tool that hands a filmmaker a clean report while missing the brand most likely to sue them does damage nobody discovers until the lawyers are already involved.

  • We stopped asking "did it work?" and started asking "can I make it fail?" Every check now needs an example that should pass and an example that should fail. If a test cannot fail, it was never testing anything, it was just making us feel better.

  • Fixing something makes you stop looking at it. We fixed a bug where acronyms like IBM were looked up incorrectly, felt good, and moved on. A fresh pair of eyes found we had only fixed half of it, and names like McDonald's were still wrong. That feeling of "handled" is not evidence.

  • Honesty turned out to be a feature. Every instinct says to hide what your product cannot do. But for a tool whose whole job is telling somebody what they might have missed, the limits are not embarrassing footnotes. They are the most useful thing on the page.

What's next for Clearance Desk

  • Disambiguate what the Wikipedia article is actually about. A character called Phoenix or Mercury matches a city or a planet. Reading the article's opening line, or its Wikidata type, would separate a real person from an unrelated concept, which is the largest remaining false-positive class.
  • Weight common names down, not up. "John Smith" scores HIGH on 32,614 pageviews; in clearance practice a very common name is safer, because no individual can claim it. This needs a threshold derived from data, not guessed.
  • Downloadable clearance report in the format production counsel actually files.
  • Suggest cleared alternatives. A flagged name is only half the job; re-running candidates through the same pipeline until one comes back CLEAR closes the loop.
  • A dedicated ClickHouse Cloud cluster, removing the shared-quota dependency and the per-scan element cap.

Built With

Share this project:

Updates

Submission history