Take2 — a second pass on your script before it's locked

Inspiration

Before a single frame of a film or TV show is shot, someone has to clear the script. Every real song on the radio, every branded can on the table, every celebrity face on a poster, and every famous landmark out the window needs licensing, clearance, or a permit. Today this is slow, manual, line-by-line work — and the one risk a tired reviewer misses is the one that becomes a reshoot, a lawsuit, or a finished film that can't be distributed.

Most AI-in-film projects reach for the glamorous, creative side — generating scripts, storyboards, or trailers. We went the other way, at the unglamorous but genuinely valuable operational work a production actually runs on. That's exactly the "execution, not experimentation" phase Google Cloud is pointing at for media and entertainment.

What it does

Take2 is an agentic rights-and-clearance system for scripts. You give it a scene, and four specialist AI agents read it in parallel:

  • Music Licensing — real songs, artists, and bands
  • Trademark & Brands — real commercial products and companies
  • Likeness & People — real public figures whose image appears on screen
  • Location Clearance — real landmarks and locations needing permits

A governing layer then reconciles every finding into a single verdict per scene — Cleared, Flagged, or Blocked — with plain-language reasoning for every flag. And every decision streams to Grafana, so the whole review is observable and auditable, not a black box.

How we built it

  • Google Cloud & Gemini via the Agent Development Kit (ADK): each domain agent is a Gemini-powered LlmAgent. In the ADK web interface, a root orchestrator delegates the scene to the four specialists — you can watch the delegation happen live in the trace and graph views.
  • Deterministic governance: in the production pipeline, the four agents run concurrently with asyncio.gather, and a plain-Python governing layer — not an LLM — reconciles their flags into a verdict. This keeps clearance decisions predictable, testable, and auditable.
  • Docling parses raw scripts and splits them into scenes.
  • Grafana Cloud is the observability layer: every verdict and flag is emitted over OpenTelemetry as metrics and logs, feeding a live "Clearance Control Room" dashboard — verdict breakdown, flags by agent, blocked rate, and a scene-by-scene audit trail. This is observability of the AI's decisions, not just infrastructure.
  • FastAPI + Vertex AI power the deployed web app, which both displays the live Grafana data and runs the full analysis pipeline on demand. It's deployed and public.

Challenges we ran into

  • Agent scoping. Early on, agents flagged into each other's domains — the trademark agent catching songs, the likeness agent catching band names. We tightened each agent's instructions with explicit exclusion rules so every specialist stays in its lane.
  • Deterministic vs. agentic orchestration. We deliberately kept the production reconciliation deterministic for auditability, while building a separate ADK-web-driven orchestrator to visualize the multi-agent delegation — two views of the same agents, each optimized for a different goal.
  • Observability plumbing. Getting OTLP metrics and logs flowing cleanly into Grafana — and querying them back through the datasource proxy for the live dashboard — took real iteration around counter naming, label casing, and query windows.
  • Making it genuinely deployable. Running a heavy ML pipeline (Docling) behind a real web server, on Vertex AI, with proper service-account auth in a container — so the whole thing is live and public, not just a local demo.

What we learned

That the strongest agentic systems pair genuine multi-agent reasoning with a deterministic, auditable governing layer — and that observability isn't a nice-to-have for an AI that makes consequential calls; it's the thing that makes it trustworthy. An AI clearing scripts for legal risk has to show its work, and Grafana is where it does.

What's next

Multimodal input — checking storyboards, posters, and set photos for visible brands, faces, and landmarks using Gemini's multimodal capabilities. Multi-studio tenancy, using the studio label already flowing through every metric. And, longer term, extending the same observable, agentic pattern across a production's entire rights surface — script to footage to marketing.

Catch it in the script, not in the reshoot.

Built With

Share this project:

Updates

Submission history