Inspiration

Test screenings are still one of the main ways a film gets evaluated before release, and the feedback they produce is almost unusable. You get adjectives. "It dragged." "I lost interest in the middle." An editor has to guess which of the several hundred cuts in a scene caused it.

The data to answer that question properly already exists, or could. People watch on players that know exactly where they paused, rewound, or stopped. What was missing was somewhere to put millions of millisecond-stamped events and query them fast enough to actually investigate, and something to do the investigating.

What it does

MomentLab records consented, second-by-second audience behavior during a screening and turns it into a specific edit worth testing.

A screening session registers with consent recorded first. Playback telemetry streams in behind that consent. A server-side detector finds where retention falls off, and a Google ADK agent running Gemini 2.5 Pro investigates by querying ClickHouse through the official ClickHouse MCP server.

In the demo experiment it finds retention dropping 18.1% at 00:37, inside a window from 00:33 to 00:41, at 95% calibrated confidence, computed across roughly 2.08 million playback events from 35,001 respondents. It proposes cutting the scene between 33 and 41 seconds. A reviewer with a token approves that, the A/B test runs, and the result comes back as a +11.1% engagement lift with a 95% confidence interval of +10.9% to +11.4%, across 17,500 respondents per arm.

Two things matter about how that is presented. Every figure carries its state, so PREDICTED means a forecast and MEASURED means a result, and a hypothesis is marked GROUNDED only if real ClickHouse queries back it. And every query the agent ran is inspectable: the SQL, the duration, the rows read. You can check the number instead of trusting it.

How we built it

The agent is Google ADK. The Agent and Runner are imported and invoked at runtime in backend/agents/mcp_client.py, and a CI hook called check-ai-compliance.py fails the build if that stops being true. The model is Gemini 2.5 Pro reached only through Vertex AI, with all four HarmCategory entries set to BLOCK_MEDIUM_AND_ABOVE.

ClickHouse is reached through the official ClickHouse/mcp-clickhouse server over stdio, against ClickHouse Cloud running 26.4.1.2212.

The part that made the interaction feel live is the aggregation. Retention lives in an AggregatingMergeTree materialized view keyed by (project, experiment, scene, respondent_cohort, media_time_ms), joining playback events to screening sessions. On a representative load that view answers in about 48ms where the same question against raw events takes around 128ms. That headroom is the difference between an agent that investigates while you watch and one that runs overnight.

Backend is FastAPI with Pydantic, frontend is React, Vite and TypeScript, both on Cloud Run with Terraform for the infrastructure and Secret Manager for credentials. Firestore holds the hypothesis documents. Scene assets were generated out of band with Veo and are labeled SYNTHETIC everywhere they appear.

Challenges we ran into

Averaging averages. The first cohort implementation computed each cohort's mean and then averaged those means for the "All" line, which is wrong whenever cohorts differ in size. The fix was to merge the underlying avg aggregate states across cohort rows, which gives the true sample-weighted mean. The number moved.

A retention cliff that was not real. The last few buckets of the timeline held a single respondent each and plotted at 0%, drawing a dramatic drop at the end of the scene that was five stray events out of two million. We added a minimum bucket sample of 10, matching the floor the product already advertises on its consent screen.

Provenance that reported nothing. The query detail modal kept showing 0ms and 0 rows. system.query_log holds two rows per query, QueryStart and QueryFinish, and the lookup was matching whichever came back first. QueryStart always has zeroes.

Agent runs that took minutes to appear. We first read recent agent activity out of system.query_log, which on ClickHouse Cloud is per-replica, so a run could be invisible for several minutes even with clusterAllReplicas. Rather than fight the visibility, we write a small momentlab.agent_runs ledger at the end of each run and read that instead.

An indicator that was lying. The screening page carried an "AGENTS ONLINE" badge showing a hardcoded 4. It was removed rather than wired up, because a research participant has no use for agent telemetry and an invented number has no business sitting next to a privacy guarantee.

Accomplishments that we're proud of

The numbers on screen are reconstructable. The dataset is seeded and the aggregates rebuild from raw events using scripts in the repo, so nothing has to be taken on trust.

The labeling discipline held up. "Verified" appears only where the system actually confirmed something, never on a forecast. When we found a status indicator that did not reflect reality, we deleted it.

The mobile work is real rather than a collapsed desktop grid. All eight primary screens were checked at 375px with no horizontal page overflow, verified by sweeping the DOM for elements extending past the viewport rather than by looking at screenshots.

What we learned

Reading a diff is not verification. Several defects here looked correct in code and only appeared when the thing was actually run: the provenance modal, the false cliff, the averaging error. Every one was caught by exercising the path, not by review.

ClickHouse rewards designing the aggregation around the question. Cohort-keyed aggregate states let one table serve the overall curve and four cohort curves without rescanning raw events, and that single decision is what makes the agent usable interactively.

What's next for MomentLab

Live respondents at scale, since the screening flow already accepts them and only the demo dataset is seeded. Completing WCAG 2.2 AA, where the known gaps are dialog focus trapping, skip navigation, and a full screen-reader and 200% zoom pass. And multi-scene experiments, so an editor can compare cuts across a whole reel rather than one scene at a time.

Built With

Share this project:

Updates

Submission history