Inspiration
Agents can read almost anything, but they can't hear a video. Every workflow that touches YouTube, TikTok, or a recording hits the same wall: captions are missing, auto-generated, or wrong, and "transcribe this" becomes a manual detour. I wanted one small, honest tool that turns any video URL or file into a clean transcript and exposes that single capability everywhere an agent or a human might need it, including directly on a web page through WebMCP.
What it does
Paste a YouTube, TikTok, Facebook, X, or Reddit link, or drop a video or audio file, and get a clean, timestamped transcript as Markdown, plain text, SRT, or JSON. It runs Whisper on the actual audio, so it works on videos with no captions at all.
One core, four surfaces:
- CLI and npm library:
npm i -g transcriptly, thentranscriptly "<url>". Whisper runs locally on your machine; the first run checks your tools and lets you pick a model with arrow keys. - MCP server:
claude mcp add --scope user transcriptly -- npx -y transcriptly mcpgives Claude Code, Cursor, Codex, ChatGPT, Copilot, or any MCP client two tools:get_transcriptandget_video_info. - Hosted API at transcriptly.dev: a free tier with honest limits, live progress streamed as server-sent events, and curl, TypeScript, and Python examples in the docs. A waitlist gauges demand for a paid, generous tier.
- WebMCP: the landing page registers
get_transcript,get_video_info, andtranscribe_filethroughdocument.modelContext. A browser agent can transcribe a video right on the page, and even transcribe a file the human dropped into the page's drop zone. An agent panel shows every tool call and its progress as it happens, and the human sees the same transcript the agent got.
How I built it
TypeScript throughout. The core resolves a source with yt-dlp (or ffprobe for files and direct media links), extracts audio with ffmpeg, runs speech recognition through whisper.cpp locally or Groq's whisper-large-v3-turbo in the hosted tier, normalizes the segments, and formats the result. The CLI runs on plain Node 20 with the MCP SDK as its only runtime dependency. The MCP server uses the official SDK over stdio. The API is a small Bun server with a transcript cache, sliding-window rate limits, a daily transcription budget, and a single shared pipeline that serves both plain responses and progress streams. The site is Next.js with a static export, hand-built animated SVG illustrations, light and dark themes, deployed with Caddy on a Hetzner box through GitHub Actions.
The WebMCP layer is deliberately thin: the page's tools call the same API the human uses, so the agent path and the human path share one implementation and one cache. The transcribe_file tool exists because of a WebMCP-specific idea: the human drops a private file into the page, and the agent gets to work with something that never had a URL.
Challenges
- YouTube, Facebook, X, and Reddit block or login-wall datacenter IPs. The hosted tier routes those platforms through a residential proxy and pre-warms a transcript cache; the CLI on your own machine has none of these problems.
- Instagram exposes only video-only streams to yt-dlp without a login, so it is listed honestly as unsupported rather than half-working.
- Groq caps uploads at 25 MB, which raw 16 kHz audio exceeds after 13 minutes; the engine re-encodes to compact MP3 before uploading.
- Agents don't discover things the way people do: an agent asked for "the file" tried the page's download buttons instead of calling the tool with
format: "md". The fix was better tool descriptions, which is where most WebMCP polish actually lives.
Accomplishments
Four working surfaces from one core, all live and tested end to end: the npm package, the MCP server inside Claude Code, the hosted API, and the WebMCP page with real tool calls from an agent browser. MIT licensed and documented for someone who has never seen it.
What I learned
WebMCP makes "the page is the tool" feel natural: once the tools exist, an agent uses the page the way a person would, and the same UI can show the person what the agent is doing. Most of the hard work was not the AI part but the plumbing around platforms, limits, caching, and honest error messages.
What's next
A paid hosted tier for longer videos if the waitlist says so, speaker labels, translation, and more platform-specific care where yt-dlp needs help.
Log in or sign up for Devpost to join the conversation.