## Inspiration

Before any independently financed film can get distribution, it needs Errors & Omissions insurance — and no insurer will write a policy until every referenceable element in the script has been cleared: character names that might match a real person in the same profession and city, business names, phone numbers, brands, music cues, source material rights, real people depicted. Professional clearance houses charge thousands of dollars and take days to weeks to do this research by hand. Independent filmmakers routinely skip it and discover the problem in post-production, when it's expensive or impossible to fix.

We wanted to see if an agentic pipeline — live web research plus LLM reasoning, not a static checklist — could do this research automatically, in minutes, with real citations a filmmaker could actually hand to a lawyer.

## What it does

Clearance takes a screenplay (pasted, or uploaded as .txt/.fountain/.md/.pdf) and:

  1. Extracts every clearance-relevant entity — person names, businesses, contact identifiers, brands, music cues, source material, real people depicted, and the title — using Gemini structured output, chunked by scene.
  2. Researches each one live on the web via the Parallel Search API, with query templates tailored per entity type so "Ellen Reyes" becomes "Ellen Reyes" cardiologist Chicago, not a useless bare-name search.
  3. Runs four concurrent ADK agents, one per clearance domain (Names & Businesses, Title Clearance, Music & Source Material, Real People Depicted), each scoring RED/AMBER/GREEN against a written rubric and citing the actual source for every claim.
  4. For anything RED or AMBER, proposes a cleared alternative — a replacement name/business/title that preserves the original's register (cultural origin, era, syllable rhythm) — and re-researches the candidate through the identical live pipeline before calling it clear.
  5. Streams results live to a single-page UI as each pass completes, with an annotated-script view, a filterable clearance report, and a full citation trail for every finding.

Every result comes from a live Gemini or Parallel call at request time. No mock data, no fixtures, no demo mode — if a service fails, the app says so.

## How we built it

  • Backend: Python 3.11+, FastAPI, streamed over Server-Sent Events so a 60-second research job doesn't leave the user staring at a spinner.

    You can add up to 25 tags.

  • Built with

    • Backend: Python 3.11+, FastAPI, streamed over Server-Sent Events so a 60-second research job doesn't leave the user staring at a spinner.
    • Reasoning: Google's Agent Development Kit (google-adk) for the four clearance-pass agents (structured output + tool-calling), google-genai for extraction and substitution generation, running on Gemini.
    • Research: the Parallel Search API (parallel-web SDK), with per-entity-type query templates, concurrent asyncio.gather execution, and retry-then-degrade on failure — a finding with no real supporting citation is treated as a bug, not an edge case.
    • Frontend: a hand-built single-page app — no framework — with a live pipeline stepper, stat cards, a searchable/filterable report, and an annotated script view with inline citation panels.
    • Deployment: Docker + Cloud Run, secrets in Secret Manager, built directly from source via Cloud Build.

Challenges we ran into

  • ADK's output_schema + tools combination turned out to only be reliable on the Vertex AI backend — on the Gemini Developer API it silently degraded to prose instead of JSON, tracked down by reading ADK's own capability-detection source.
  • Free-tier quota surprises: our first model choice had a 20-requests/day cap (not per-minute) that we blew through during testing; moved to a model with a 500/day quota and added proper 429-aware retry logic that respects Google's own suggested backoff.
  • Google Cloud billing rabbit hole: assumed a $1,000 "GenAI App Builder" trial credit would cover Gemini calls — it doesn't; that credit is scoped to Vertex AI Search & Conversation, a different product entirely. Ended up building the whole thing to run cleanly on the free Developer API tier instead, with Vertex as a one-env-var upgrade path.
  • A from-scratch Docker build caught dependency pins our local dev environment had silently drifted past — pydantic, google-cloud-aiplatform, and fastapi were all pinned to versions incompatible with each other and with google-genai, invisible locally because the venv had already resolved to newer versions before the requirements file was ever tested clean.

Accomplishments we're proud of

A fully live, zero-mock agentic pipeline — real Gemini reasoning, real Parallel citations, four concurrent domain-specific agents, a working clear-then-substitute loop that actually re-verifies its own suggestions — deployed and reachable at a public URL, built entirely on free-tier infrastructure.

What we learned

That "the free tier" is not one thing — Developer API vs. Vertex AI have meaningfully different capabilities (structured output + tools together, quota shapes, billing requirements), and that promotional credits are far more narrowly scoped than their names suggest.

What's next

Vertex AI for production-scale quota, OCR support for scanned/image-only PDFs, and expanding clearance coverage toward the audio/visual side of the pipeline (music licensing status via composition + master rights lookups, VFX/likeness clearance for archival footage).

  • Backend: Python 3.11+, FastAPI, streamed over Server-Sent Events so a 60-second research job doesn't leave the user staring at a spinner.
  • Reasoning: Google's Agent Development Kit (google-adk) for the four clearance-pass agents (structured output + tool-calling), google-genai for extraction and substitution generation, running on Gemini.
  • Research: the Parallel Search API (parallel-web SDK), with per-entity-type query templates, concurrent asyncio.gather execution, and retry-then-degrade on failure — a finding with no real supporting citation is treated as a bug, not an edge case.
  • Frontend: a hand-built single-page app — no framework — with a live pipeline stepper, stat cards, a searchable/filterable report, and an annotated script view with inline citation panels.
  • Deployment: Docker + Cloud Run, secrets in Secret Manager, built directly from source via Cloud Build.

Challenges we ran into

  • ADK's output_schema + tools combination turned out to only be reliable on the Vertex AI backend — on the Gemini Developer API it silently degraded to prose instead of JSON, tracked down by reading ADK's own capability-detection source.
  • Free-tier quota surprises: our first model choice had a 20-requests/day cap (not per-minute) that we blew through during testing; moved to a model with a 500/day quota and added proper 429-aware retry logic that respects Google's own suggested backoff.
  • Google Cloud billing rabbit hole: assumed a $1,000 "GenAI App Builder" trial credit would cover Gemini calls — it doesn't; that credit is scoped to Vertex AI Search & Conversation, a different product entirely. Ended up building the whole thing to run cleanly on the free Developer API tier instead, with Vertex as a one-env-var upgrade path.
  • A from-scratch Docker build caught dependency pins our local dev environment had silently drifted past — pydantic, google-cloud-aiplatform, and fastapi were all pinned to versions incompatible with each other and with google-genai, invisible locally because the venv had already resolved to newer versions before the requirements file was ever tested clean.

Accomplishments we're proud of

A fully live, zero-mock agentic pipeline — real Gemini reasoning, real Parallel citations, four concurrent domain-specific agents, a working clear-then-substitute loop that actually re-verifies its own suggestions — deployed and reachable at a public URL, built entirely on free-tier infrastructure.

What we learned

That "the free tier" is not one thing — Developer API vs. Vertex AI have meaningfully different capabilities (structured output + tools together, quota shapes, billing requirements), and that promotional credits are far more narrowly scoped than their names suggest.

What's next

Vertex AI for production-scale quota, OCR support for scanned/image-only PDFs, and expanding clearance coverage toward the audio/visual side of the pipeline (music licensing status via composition + master rights lookups, VFX/likeness clearance for archival footage).

Built With

  • css3
  • fastapi
  • gemini-api
  • google-adk-(agent-development-kit)
  • google-cloud-build
  • google-cloud-run
  • google-secret-manager
  • html5
  • javascript
  • parallel-search-api
  • pydantic
  • pypdf
  • python
  • server-sent-events
Share this project:

Updates

Submission history