OpenCast

Inspiration

Two ideas collided. The first: tools like Descript proved that the transcript is the best interface for editing talk-heavy video — you delete a sentence, the footage follows. The second: WebMCP proposes that a web page should hand agents typed, callable tools instead of making them guess at buttons.

Put together, they describe the same product. If editing is already "operate on words," then an agent that speaks in words is a native editor, not a bolted-on chatbot. I wanted to build the version of that idea where the human and the agent aren't in separate worlds — they look at the same live page, pull the same levers, and every change lands in the same undo stack. Editing a podcast felt like the sharpest possible test: it's real work, it's language-shaped, and you can see whether the collaboration is real.

What it does

OpenCast is a transcript-first podcast/video editor. Upload a recording; the transcript becomes editable in about a minute; then you — or your agent, through 47 WebMCP tools — cut filler words and silences, fix mis-transcriptions, assemble named compositions (hooks, clips) from timeline ranges, turn on live captions, layer timed images over or under the footage, remove the background with in-browser ML, and export a finished MP4 with everything burned in — rendered entirely in the browser.

How I built it

The architecture is three layers, and it's the part I'd defend in any design review:

  • Page state — a Zustand store holding words (with per-word timestamps), sources, compositions, overlays.
  • One shared action hub — every edit is a narrow verb: cutWords, renameSpeaker, addToComposition, removeOverlay…
  • Two front ends to the same hub — buttons for humans, and WebMCP tools (document.modelContext.registerTool — a name, a JSON schema, a callback) for agents. Same verbs; nobody gets a special door. Tools return what changed, so the agent can verify itself, and destructive ones require explicit confirmation.

Around that core: a Next.js app on Vercel; a Fly.io Docker worker that takes resumable 16 MiB upload chunks straight to a volume, extracts audio with ffmpeg, and runs a words-first pipeline — Whisper word timings through a bounded concurrent pool unlock the editor fast, while speaker diarization runs as a background enrichment that relabels the live transcript when it lands (and multi-track projects skip ML diarization entirely, because the track is the speaker). Exports use a canvas compositor: play the kept ranges once, draw background layer → footage (MediaPipe selfie segmentation for background removal) → B-roll layer → captions, and record it with the source audio via MediaRecorder to MP4.

Challenges I ran into

  • The silent five-minute ceiling. Long diarization calls kept dying with a cryptic Headers Timeout Error: Node's fetch (undici) aborts any request whose response headers take over 300s — and retried the doomed request forever. Diagnosing that, giving OpenAI calls a real deadline through an explicit dispatcher, and then restructuring to words-first (never let the slow model block the editable transcript) was the single biggest turn in the project.
  • Embedded browsers are a different planet. ChatGPT Desktop's browser is narrow, blocks programmatic downloads without a user gesture, and can't be trusted with a GPU delegate. That forced real fixes: on-screen layers scoped to the footage's aspect ratio instead of the pane, a CPU fallback for segmentation, and layouts audited at 480px.
  • Canvas taint. Recording a canvas fails if any drawn image isn't CORS-readable — one bad overlay URL could poison a whole render. Locally uploaded images are compressed into data URLs (always safe); non-CORS remote images are skipped rather than killing the export.
  • Building on a moving spec. WebMCP is pre-standard; API names have already shifted. Keeping the tool layer thin over the action hub meant the churn only ever touches one file.

What I learned

Design tools as narrow verbs with clear side effects that report what changed — agents chain small verbs better than they wield big ones. Words-first beats complete-but-slow: shipping the editable thing early and enriching in the background changed how the product feels more than any feature. And the deepest one: the trust problem with agents isn't solved by logs or promises — it's solved by a shared screen, where every mutation happens in front of you and your undo key still works.

What's next

Program-cut-aware multicam rendering on the worker, cross-chunk speaker identity stitching, and a consent surface for tool calls that's worthy of the spec this is all betting on.

Built With

  • codex
  • fly.io
  • openai
  • vercel
  • webmcp
Share this project:

Updates

Submission history