Inspiration

In 2020 Netflix wrote one line of dialogue: "There's Nona Gaprindashvili, but she's the female world champion and has never faced men." She had faced fifty-nine of them. She sued for $5 million, a federal judge refused to dismiss it, and Netflix settled.

One sentence, about a real person, in a script nobody checked.

It keeps happening, and the pattern is always the same. Squid Game put a real Korean phone number on a business card and the woman who owned it took over 4,000 calls. Warner Bros. told a federal court it was prepared to digitally alter The Hangover Part II to put a different tattoo on Ed Helms's face. Netflix and Tyler Perry retitled a series in the middle of its own season after a trademark suit filed three days before the premiere.

The worst version is when it surfaces late. On Michael, Lionsgate found a 1993 settlement clause after principal photography had wrapped: 22 days of additional photography, $10–15M added to a $155M budget, the release slipped a year, and the reshoots no longer qualified for the state tax rebate.

Every one of those is a name, a number or a reference that nobody checked.

What it does

Clearance reads a screenplay and produces the report an E&O insurer asks for.

It extracts every proper noun, classifies each one, searches it against the live web through Parallel, and decides whether a hit is a collision or a coincidence. Every finding cites the page it came from, because a clearance report nobody can check is an opinion.

The console has four views. Findings is what a lawyer reads first, worst first. Every item is the whole extraction, including what was dismissed and why. Sources is every page the report rests on. How it decides is the reasoning in the open.

How I built it

Google Cloud: Gemini through Vertex AI, deployed on Cloud Run with the API key in Secret Manager. Partner: the Parallel Search API through the official parallel-web SDK.

Four stages, split because they fail differently:

  • find by rule — phone numbers, addresses, plates, sluglines and character cues have a shape, so a regex finds them. It does not hallucinate and costs nothing. It also knows that 555-0100 to 555-0199 is the block reserved for fiction, and that INT. TRUCK STOP names a kind of building rather than a business. Both settle without spending a search.
  • find by reading — a company named mid-dialogue has no shape. Gemini reads two scenes at a time and lists candidates only. It is never asked to judge them, because a stage that both finds and dismisses will quietly dismiss the thing it failed to look up.
  • search — one Parallel call per item, eight concurrent.
  • adjudicate — Gemini weighs how close the match is, how distinctive the name is, and what the script does to it. If it is genuinely unsure it writes a better query and searches once more. Once, because an agent that can re-query forever will.

Challenges I ran into

Batching looked free and wasn't. Parallel's search() takes a list of queries, and twenty queries in one call returns as fast as one. But the response caps at ten results and carries no indication of which query produced which. A clearance report has to say this name collides with that company, so I measured it, threw the batching away, and went one call per item at concurrency eight. Sixteen items take 6.8s at four, 2.3s at eight, 2.3s at twelve.

Screenplays wrap, and entities straddle the wrap. "My office is at 233" ends one line of dialogue and "South Wacker Drive, Chicago." begins the next. Extracting line by line found neither half, so a real address was silently missing from the report. Lines are now merged into blocks before extraction.

The cue is not the name. I resolved every character's full name and then never used it, so the agent searched "HALLORAN" and returned a character from The Shining. Searching "Marcus Halloran" returns people who could actually object.

One company, three entries. The script says "Northgate Pharmaceuticals", then "the Northgate plant", then "Northgate". The model reads two scenes at a time, so no single call sees all three. Folding them is deterministic now, and it treats a generic premises word as noise, so the same company reaches a lawyer's desk once instead of three times.

The adjudicator was unstable on the borderline cases. TRUCK STOP came back none on one run with a good argument, and medium on the next citing truckstop.com, which is logistics software rather than a physical truck stop. A report is only useful if its noise level is stable, so generic slugline locations are now settled by rule instead of re-argued every time.

The research contradicted my own pitch. I had been saying clearance is expensive. It is not: about $1,000 for a feature, and that price has not moved in ten years. I dropped the claim. I had also been implying a knowable item count, and no vendor, sample report or court exhibit publishes one, so I measured my own fixtures instead and label it as mine.

Accomplishments that I'm proud of

I wrote a screenplay to test it and invented a company called Northgate Pharmaceuticals, because it sounded made up, and gave it a recalled-insulin scandal. The first search returned northgatepharma.com, a real pharmaceutical distributor.

I had put a real company in a film about poisoning diabetics, and I never checked, because a name that sounds invented is exactly the one nobody checks. That is the product demonstrating itself on its own author.

The other thing I am pleased with is what it doesn't flag. Three items come back with no collision, and two are settled by rule before a search is spent. A report that flags everything costs a lawyer the same reading time as no report at all.

What I learned

That the hard part was never the lookup. Extraction is a regex, search is an API call. Almost every string returns results, so the entire product is the judgement about which of them mean anything.

And that a research pass worth doing is one that can prove you wrong. I asked for the clearance-cost research adversarially and it took two claims off me.

What's next for Clearance

Read the PDF and Final Draft formats a production actually circulates, not just Fountain. Track a script across revisions so only new names are re-checked, the way vendors charge $10 per new name. And export the report in the shape an E&O application expects, since that is the document it exists to satisfy.

Built With

Share this project:

Updates

Submission history