Inspiration

An editor can watch a shot dozens of times. An audience usually sees it once. That gap matters: what is obvious to the person making a film may not be what a viewer remembers.

Recall Cut asks a practical question before another edit: what did viewers actually report, and does that evidence justify changing this moment?

What it does

Recall Cut connects audience recall to an editor's decision. The participant flow separates consent, a single playback, a 20-second distractor task, free recall, and a follow-up continuation. Responses are frozen before analysis, and protocol deviations remain visible.

The creator records an initial editorial hypothesis before reviewing audience evidence. Gemini extracts structured claims linked to exact quotes and character offsets. Cohort analytics distinguish matching details, explicit disagreement, uncertainty, and details not mentioned. Counts represent distinct eligible viewer sessions, not the number of claims an AI extracted.

The editor can inspect the evidence, develop a time-specific proposal, revise the rationale, or choose Keep Original. Decisions and revisions are saved. An omission is not treated as proof of misunderstanding, and a proposal is a hypothesis to test—not an automatically proven improvement.

How we built it

The application combines a Next.js and TypeScript creator workbench with a Python/FastAPI backend and PostgreSQL persistence. Google Cloud Run hosts the application, with Cloud SQL, service-to-service authentication, and Secret Manager supporting deployment.

The live analysis path uses Gemini through Google's model API and the Google ADK runtime. A dedicated ClickHouse writer ingests session and claim records. The ADK analyst retrieves analytical results through the official mcp-clickhouse server using a separate read-only account. ClickHouse is part of the evidence workflow, not merely a logging destination.

Challenges we ran into

Counting extracted claims initially overstated how many viewers remembered a detail. We changed the analytics to distinct-session counts and exact cross-phase unions, including eligible viewers who produced no claims. We also addressed stale record versions, durable persistence, query-result validation, and the distinction between unavailable data and a genuine zero.

Another challenge was making a small agent workflow predictable: model request limits, explicit live-execution gates, no silent fallback to a mock result, and saved reports that can be reopened without generating them again.

What we learned

A small formative exercise with three viewer reports and an editor helped sharpen the product question. After reviewing those reports, the editor shifted from a hypothesis about giving a couple's interaction more time to a narrower hypothesis about clarifying the tablecloth pickup. This was qualitative discovery, not a controlled study or evidence that an edit improved comprehension.

The useful outcome is not always another cut. Sometimes the evidence supports restraint.

Current verification boundary

A bounded synthetic integration test completed through the deployed application, with a saved live-mode report recording five model requests, three official ClickHouse MCP queries, five validated claims, and successful study-scoped reconciliation. The saved result was independently reopened through the hosted API. This verifies a synthetic integration path, not Coffee Run source accuracy or an improvement in viewer comprehension. The acceptance fixture used invented scene details and unverified anchors; its model-boundary telemetry label also requires correction. The public preview was restored to demo mode after testing. Judge-facing live release verification remains outstanding. Video generation and automatic rendering are disabled and are not claimed as delivered features. Synthetic acceptance fixtures remain separate from human viewer evidence.

What's next

Complete hosted live acceptance, run independent comparisons of an original shot and an editor-approved timing variant, and expand beyond the current benchmark workflow. Any claim that a variant works better must come from new viewer evidence.

Sources and attribution

The research benchmark uses a shot from Blender Studio's Coffee Run: https://studio.blender.org/projects/coffee-run/?asset=3300. The film is third-party source material, not our creation. The interface adapts the Toluva template: https://github.com/Pavilion-devs/toluva. Asset licenses and attribution must be retained and checked against the submission rules.

Built With

Share this project:

Updates

Submission history