Inspiration

Every creator with any traction has stared at their YouTube retention graph, watched it dip at some random timestamp, and had zero idea why. YouTube Studio shows you that people left — never why. The obvious hackathon idea for "automate the creator workflow" is another caption or clip generator. We wanted to build something that diagnoses a problem creators already have, using data they already own, instead of generating more content for them to manage.

What it does

You upload your video's transcript and your real YouTube Studio retention CSV export. Retention Detective finds the moments where your retention dropped anomalously — not just where it's low, since every video naturally declines, but where the drop is statistically steeper than that specific video's own normal decay rate. For each anomaly, an AI investigates like an analyst would: it pulls the surrounding transcript, checks your speaking pace, compares the drop's severity against your video's own baseline, and checks whether it's part of a recurring pattern — before reaching a verdict. Every diagnosis shows its full evidence trail, so you can see exactly why the AI concluded what it did, not just take its word for it.

How we built it

  • Detection engine: a median-baseline decay algorithm (Node/TypeScript) that flags real anomalies without false-positiving on normal video decay.
  • Investigation loop: an agentic tool-calling pipeline (Featherless, Qwen3-32B) with five real tools — transcript window lookup, retention shape classification, baseline comparison, pacing analysis, recurring-pattern search. The model gathers evidence before concluding, it doesn't guess from one snippet.
  • Transcription: Groq Whisper for real word-level timestamps when a creator uploads audio/video instead of pasting a transcript.
  • Frontend: Next.js, wired directly to real analysis data end-to-end — no mock data anywhere in the shipped app.
  • 82 automated tests covering the parser, the detection algorithm, every tool, and the full investigation loop.

Challenges we ran into

Our first version of the anomaly detector used the whole video's average decay rate as its baseline — and a big enough anomaly would inflate that very baseline, hiding itself from its own comparison. We caught this with a test proving "a smooth, healthy video should flag zero anomalies," fixed it by switching to a median of per-segment decay rates (robust to a small number of outliers), and re-verified against a deliberately constructed test case. We also caught the model narrating conclusions about tools it never actually called ("pacing analysis would reveal...") and tightened the system prompt so every claim in a diagnosis has to trace back to a tool call that actually ran.

Accomplishments that we're proud of

Getting the investigation loop to behave like a real detective instead of a chatbot — gathering evidence, sometimes ruling things out (one live test correctly found normal pacing and pointed to weak content instead), and showing its work. That auditability was the whole point, and watching it work live for the first time, with a diagnosis that correctly traced back to four separate pieces of real evidence, was the moment we knew the idea actually held up.

What we learned

That the hardest part of "AI diagnoses your data" isn't the AI call — it's building the deterministic groundwork (parsing, baseline math, tool design) rigorous enough that the AI has something honest to reason over. Bad math in, confident-sounding nonsense out.

What's next for Retention Detective

  • "Cross-Examine This Finding" — already built — lets a creator ask follow-up questions about a diagnosis, answered strictly from the evidence already gathered, not a fresh guess.
  • Next: a "what if" simulator that projects retention impact if a flagged segment were edited to match the video's own healthy pacing, and direct YouTube Analytics API integration so creators skip the manual CSV export entirely.

Built With

Share this project:

Updates

Submission history