Inspiration
Music clearance is one of the most common rights problems in documentary and independent film. The free tools mostly do one of two things:
- Fingerprint audio to estimate whether a platform will flag it
- Check a static public-domain list that usually covers compositions rather than specific recordings
One licensing service's own guidance is essentially to look up the publication year and, if it's before 1931, treat the song as clear.
That's a filter, not an answer. It clears a small set of very old songs and says nothing about everything after 1930, which is most of what a filmmaker might actually want to use. It also addresses only half the basic problem, because a song isn't one copyright. It's two.
The composition is the song itself, music and lyrics, associated with its writers and publishers. The sound recording is a particular recorded performance, controlled by whoever owns the master. Those rights follow different rules and can expire decades apart.
Take Louis Armstrong's 1928 recording of "West End Blues." In the US, the composition entered the public domain on January 1, 2024, but the recording is protected until 2029 under the Music Modernization Act. You can perform the song, but you can't use that recording.
Switch to the UK and the blocking layer swaps. The recording's term expired in 1979 and the 2013 extension didn't revive it, but the composition remains protected based on the life of its last surviving writer.
Same song, and both answers are no, for opposite reasons.
What it does
Type in a song and artist and tell it what you're making. The tool returns a cited rights assessment showing:
- Whether the requested use appears clear
- Separate verdicts for the composition and recording
- The specific layer blocking the use
- Who appears to control the relevant rights
- An estimated licensing range where one can reasonably be given
- The evidence behind each fact
- Anything the system couldn't verify
Every record says plainly that this is research, not legal advice.
Territory and intended use are controls because both can change the answer. Re-recording the song yourself drops the sound-recording layer entirely. Duration can change what a license costs without changing whether permission is required.
If the system can't establish which recording you mean, it stops and asks. A title and artist can still lead to:
- Original masters
- Reissues and remasters
- Live recordings
- Alternate takes
- Other separately controlled recordings
Guessing the wrong recording means researching the wrong copyright.
You can also paste in an entire cue sheet, and the same analysis runs on each line with the most restrictive results surfaced first. When permission is necessary, the tool drafts the sync and master-use requests from the record, leaving production-specific details as marked blanks rather than inventing them.
How I built it
Research runs through three tiers, cheapest first.
Tier 1: Licenses that settle the question directly
Tier 1 checks Creative Commons license relations on MusicBrainz against a static license table, with no model involved. A qualifying license on the work settles the composition layer, and one on the recording settles the recording layer. Release-level licenses are treated more cautiously because a license on one release doesn't establish the recording's status everywhere it appears. If Tier 1 settles a layer, nothing further is spent researching it.
Tier 2: Structured facts
Tier 2 queries structured sources directly: MusicBrainz for works, recordings, releases, writers, and the composition/recording split, and Wikidata for dates and writer death years.
These sources are cheap and useful, but the bugs I found taught me not to confuse structured data with established fact.
Tier 3: Research
Tier 3 uses Parallel for two different jobs. Search looks for what the structured databases don't reliably contain: renewal evidence, original publication and release dates, writer identities and death years, and corroboration when sources disagree.
A separate task researches current rights holders, including publishers, administrators, labels, and ownership shares. Each field carries its own citations rather than inheriting confidence from the copyright calculation. Enrichment runs after the verdict returns, so a slow publisher lookup never delays an otherwise complete answer.
The consistency layer
Between research and the rules engine sits a consistency layer that checks facts against the other facts constraining them:
- A recording shouldn't predate the composition it records
- A writer shouldn't die before a work they're credited with was published
- A writer's lifespan shouldn't be implausible for the work
- Ownership shares shouldn't total more than 100%
- A reissue date shouldn't silently become the original publication or recording date
- A life-plus calculation shouldn't run until the writer list is sufficiently established
When facts conflict, the system degrades confidence, identifies the conflict, and opens a question instead of arbitrarily trusting one side.
The rules engine
Copyright terms are computed by hand-written deterministic rules. The engine covers the Music Modernization Act schedule, life-plus terms, publication-based terms, renewal requirements, and territory-specific calculations, all as ordinary functions with known inputs and testable outputs.
No model calculates a copyright term. If a model did the arithmetic, I couldn't tell you why any answer was right.
Where Gemini is used
Gemini has one narrow job: read search evidence into a cited fact, or abstain. A finding carries supporting evidence, a citation, and a confidence level, or it becomes an open question instead of a fact.
The model interprets evidence. It doesn't invent the evidence or do the copyright arithmetic.
Infrastructure
Google ADK runs the agent graph with deterministic stages as function nodes, served by FastAPI on Cloud Run. Pipeline events stream during each run, and a persistent run log stays with every record. That means a cached result can return in under a second without becoming a black box: the record still shows what originally ran, which tier answered, and what came from cache.
Challenges I ran into
Finding the right recording
Recording selection turned out to be harder than much of the rights research. A popular song can have dozens of MusicBrainz entities:
- Original releases
- Reissues and remasters
- Compilations
- Regional editions
- Duplicates
Many aren't linked cleanly to the underlying work. My first approach reconstructed the recording history with roughly ten calls and took 71 seconds. Replacing that sweep with one work-centered search cut the selection step to about 7.5 seconds. A full cold run with Tier 3 renewal research lands at 30 to 40 seconds, and caching brings warm queries under a second.
External-service failures
MusicBrainz returned 503s on roughly one in four first attempts on some test runs, so Tier 2 isn't allowed to take the analysis down with it. A failed structured-data call is recorded and the pipeline degrades to the next research tier.
I also hit intermittent TLS handshake failures from Cloud Run to MusicBrainz. The same deployment worked one day and failed the next depending on the outbound path through Google's shared egress pool. Routing traffic through Cloud NAT with a static IP eliminated the problem.
The renewal window
The hardest remaining problem isn't a model or infrastructure problem. It's access to the evidence.
US works published from 1931 through 1963 had to be renewed after their initial 28-year term to keep protection, and those renewal records are the hardest thing in the system to reach. Pre-1978 renewals live largely in scanned volumes of the Catalog of Copyright Entries, where search often finds the correct volume without finding the actual line. Later records sit in Copyright Office systems that ordinary web search can't inspect at the individual-record level.
Accomplishments I'm proud of
The system declines to guess
The abstention that matters runs in one direction. Across every live renewal case to date, the system has never concluded a work was free without the required primary evidence.
Evidence pointing toward protection is accepted more readily. A publisher's own notice, for example, resolved "Blue Moon" as renewed at medium confidence, while weaker evidence can lean a verdict toward protected and leave the underlying question open on the record.
Declining to call a work free without sufficient evidence was the behavior I was testing for.
Evidence requirements are intentionally asymmetric
The two mistakes have different consequences. A false "protected" result may send someone after a license they didn't need. A false "public domain" result may put copyrighted material into a finished film.
So the validator treats them differently. Evidence can support or lean toward protection at lower confidence, while clearing a renewal-dependent work requires primary evidence strong enough to support that conclusion.
That rule lives in a validator, not a prompt.
Cost follows uncertainty
The research pipeline is designed so the expensive work only happens when the cheaper sources can't answer the question. A Creative Commons license can settle a layer without a model call, structured data handles the next pass, and Parallel research is reserved for facts that still need investigation. Cached results skip the research path entirely.
That makes cost a consequence of how difficult a rights question actually is, rather than a fixed cost paid on every query.
Open questions are actionable
When the evidence isn't enough, the system doesn't stop at unknown. It hands over the unresolved question, the record system most likely to answer it, and the search terms needed to continue. You can bring the evidence back and re-run the analysis.
- A bare assertion stays low confidence and can never clear a work
- An attested source can earn medium confidence
- Primary evidence can settle what secondary evidence can't
What I learned
Several early bugs looked unrelated, but each had the same underlying failure: a plausible answer built on a fact that had never actually been established.
- MusicBrainz's
first-release-datereturned 1975 for a recording from a 1928 session. It was actually the earliest release represented in the available data. Used blindly, that's a 47-year error on one of the inputs the calculation depends on - MusicBrainz credited "West End Blues" to King Oliver but omitted Clarence Williams. Life-plus-70 depends on the last surviving author. Oliver died in 1938 and Williams in 1965, so the missing writer moves the calculation 27 years toward an earlier public-domain determination
- A Wikidata search for Clarence Williams returned the actor who died in 2021, not the songwriter who died in 1965. Nothing about the result looked broken. It was a real person with a real death date from a real database, attached to the wrong identity
Two more followed the same pattern: an author whose death date preceded the work's supposed publication, and a recording dated before its own composition. In each case, the individual facts looked plausible. The relationship between them exposed the problem.
Those five failures became the consistency layer. Instead of patching each bad result, I looked for what had to be true for the whole class of result to be trustworthy: each failure became a failure class, the class became an invariant, and the invariant became a validator.
The larger lesson was that citation isn't enough. A source can be perfectly real and still refer to the wrong person, the wrong release, the wrong recording, or the wrong date. That's why every material fact carries its sources and confidence level, or isn't allowed to behave like a fact.
What's next
Two integrations would make the biggest difference in the unresolved renewal cases:
- Direct Copyright Office catalog integration — makes post-1978 records accessible to the research pipeline instead of unreachable through ordinary web search
- Reading the actual text of Catalog of Copyright Entries volumes — makes older renewal evidence searchable at the line level the system needs
Together, they could turn many of today's abstentions into evidence-backed answers without lowering the standard required to reach them.
MLC API access is also pending. Once available, registry ownership data can supersede researched rights-holder information wherever the registry provides the stronger record.
The clearance worksheet is also a step short of the document a production actually files. Adding timings and use types to the same cue-sheet pipeline would let it produce a filing-ready cue sheet for submission to the relevant performing rights organizations.
Text and film are the next asset types I want to add. For books, HathiTrust and Copyright Office records can provide some of the same underlying publication and rights evidence that MusicBrainz provides for music, while the same layer model can separate the work from particular editions and translations. Film and archival media add more layers, but the underlying approach is the same.
Music is where I kept the initial scope because it's where the problem concentrates for documentary filmmakers and where enough of the public record exists to support a citable answer.
Built With
- cloud-run
- fastapi
- firestore
- gemini
- google-adk
- google-cloud
- httpx
- musicbrainz
- parallel
- pydantic
- pytest
- python
- react
- sqlite
- tailwind
- vertex-ai
- wikidata


Log in or sign up for Devpost to join the conversation.