Inspiration

A film can have subtitles and still lose part of its story. A phone rings off-screen, a speaker is never identified, or dialogue flashes past too quickly to read. I built FrameKind to give editors a practical review room for those gaps.

What it does

FrameKind compares a short film with its existing captions. It flags timing and readability issues, missing dialogue, sound captions, and speaker context. An editor can jump to the evidence, approve or reject a proposed repair, then export SRT, WebVTT, JSON, or an HTML review report. The original captions remain unchanged.

The instant sample uses authored reference annotations and an original illustrated scene with Google-generated voices. Custom live reviews analyze the uploaded media. Caption-only reviews clearly leave audio and visual dimensions unassessed.

How I built it

The app uses React and Vite for the review workspace, FastAPI for uploads and progress events, and SQLAlchemy for run history and saved decisions. An eight-step Google ADK workflow coordinates ingestion, transcription, sound detection, visual review, timed-text rules, alignment, semantic checks, and repair planning. Gemini 2.5 Flash and Pro run on Google Cloud Vertex AI. Small media files can travel inline, with Cloud Storage supported as an alternative.

The first version was built with Google Antigravity. Replit Agent helped inspect the repository and verify its runtime and database setup. The app is publicly hosted on Replit with Postgres persistence. The repository records development tools separately from runtime services.

Challenges

The hard part was making each proposed repair traceable and safe to apply. Timing changes and text changes must compose without mutating the original file. Repeated exports must agree. Real multimodal results also exposed cases that the authored sample did not cover, including one audio-description cue overlapping multiple speech segments.

What I learned

A useful review agent needs boundaries as well as model intelligence. Evidence can be incomplete, thresholds depend on the delivery profile and the editor should retain the final decision. FrameKind treats its scores as review aids, not certification.

What's next

I want to test FrameKind with working caption editors and Deaf and hard-of-hearing viewers, improve speaker labeling and add configurable delivery profiles. Longer films will need a durable job queue and a review workflow that works scene by scene.

Built With

Share this project:

Updates

Submission history