Inspiration

At Nerdearla — Latin America's biggest developer conference — thousands of attendees sit through talks delivered in English and Spanish. Anyone who isn't fluent in the speaker's language loses half the content. Professional simultaneous interpretation costs thousands of dollars per room, and existing captioning tools require attendees to install apps, log in, and accept terms before they can read a single word.

We built Live Subtitles to make multilingual conferences effortless: one browser tab for the speaker, one QR scan for the audience, real-time captions in both languages — no install, no login, no friction.

What it does

Live mode — A speaker opens the web app, clicks "Start broadcasting," and their microphone audio streams over WebSocket to a FastAPI backend. The pipeline transcribes with faster-whisper (local, CPU, int8) and translates English↔Spanish with DeepL. Audience members scan a QR code and read bilingual captions on their own phones within ~2 seconds of the speaker finishing a sentence.

Upload mode — Drop in a recording of a past talk (mp4, mov, mp3, wav, anything). The backend transcribes it in the background and produces downloadable .srt, .vtt, .json, and .txt files. Perfect for post-production subtitling, YouTube uploads, or archiving the conference.

Multi-session — A keynote and three breakout rooms can run simultaneously. Each session has its own audio buffer, translation queue, and listener group. Sessions persist to Firebase Firestore so transcripts survive restarts and are queryable from any device.

Exports — Every session produces standard caption files that drop straight into YouTube Studio, Premiere, DaVinci Resolve, or VLC.

How we built it

Backend (Python 3.11 / FastAPI)

  • app/main.py — REST endpoints for session CRUD, file uploads, and subtitle exports; WebSocket handlers for audio ingest (/ws/audio/{id}) and caption broadcast (/ws/subtitles/{id})
  • app/session_manager.py — one ConferenceSession per talk, each its own asyncio task. Live mode uses VAD-gated buffering with partial/final segmentation. Upload mode decouples transcription from translation via an asyncio queue, then translates in parallel batches of 25 so a 30-minute talk finishes in ~90 seconds instead of 10 minutes
  • app/transcription.py — faster-whisper wrapper with domain-vocabulary biasing (initial_prompt includes FastAPI, WebSocket, CTranslate2, etc.)
  • app/translation.py — DeepL REST client using requests in a thread executor, no extra SDK needed
  • app/firebase_client.py — Firebase Admin SDK bootstrap; Firestore writes are best-effort and never block the live caption pipeline

Frontend (single-file HTML + Tailwind CDN + vanilla JS)

  • Speaker console with live audio meter, QR code generation, and share link
  • Listener view with two-column history (original + translation), font-size controls, language filter tabs, and download buttons
  • AudioWorklet captures raw PCM16 at 16 kHz — no MediaRecorder, no container parsing, no ffmpeg dependency on the client

AI

  • Whisper small model, int8 quantized, running fully local on CPU. No audio ever leaves the machine except for translation text
  • DeepL for high-quality EN↔ES (and any other language pair DeepL supports)

Challenges we ran into

The MediaRecorder trap. Our first version sent WebM/Opus chunks from the browser. faster-whisper couldn't parse them because only the first chunk contains a container header. We rewrote the capture path to use AudioWorklet, sending raw Int16 PCM at 16 kHz — decoded server-side with np.frombuffer. That fixed transcription permanently.

HuggingFace lock file timeouts. Every taskkill /F left a stale lock in the HF cache, and the next startup waited ~3 minutes for a timeout. We added local_files_only=True and a lock-clearing helper, cutting startup from 2:50 to 4 seconds.

Race condition in translation display. During uploads, all final messages fire in ~1 second, then DeepL returns 25 translations in one batch — arriving out of order. We added per-segment index fields and a pendingTranslations buffer on the frontend so translations can arrive before their matching rows and still land correctly.

Translation latency. Google Translate's free endpoint throttles at 5 req/s. Switching to DeepL with batched requests brought a 35-second sample from ~30 seconds of translation tail down to ~3 seconds.

Accomplishments we're proud of

  • Sub-2-second latency from speaker pause to visible translated caption
  • Zero client install — the audience just scans a QR code
  • Fully offline-capable transcription (Whisper runs locally)
  • Multi-session architecture that scales cleanly from 1 talk to a full conference
  • One-click SRT/VTT export from any session

What we learned

  • WebSocket streaming for audio requires raw PCM, not containerized chunks
  • Batching translations in a single HTTP request beats N parallel calls by an order of magnitude
  • A single-file HTML app with Tailwind CDN is a completely viable hackathon frontend
  • Local Whisper inference is fast enough for real-time captions when you quantize correctly

What's next

  • True incremental streaming ASR using LocalAgreement-2 for word-by-word captions
  • QR-code-based session joining with optional Firebase Authentication for password-protected talks
  • Firestore real-time listeners as a redundant caption channel
  • GPU inference on a shared worker pool for 50+ simultaneous sessions
  • On-device translation for fully offline deployment

Built With

Share this project:

Updates

Submission history