💡 Inspiration

Access to information should be effortless, accessible, and private. Millions of students, visually impaired individuals, and multitaskers struggle to digest long-form text—whether it is scanned academic papers, PDF manuals, or eBooks. Existing text-to-speech tools often locked natural-sounding voices behind expensive subscriptions, failed when reading scanned images without text layers, or uploaded sensitive private documents to remote cloud servers.

We built Blob to eliminate these barriers: a 100% client-side, privacy-first document reader and audiobook creator that transforms complex documents into natural-sounding neural audio across 140+ languages—completely inside the user's browser.


🚀 What It Does

Blob converts text, PDFs, Word files, presentation decks, and scanned images into continuous high-fidelity audio:

  • Multi-Format Ingestion: Instant parsing of .pdf, .epub, .docx, .pptx, and .txt files.
  • Scanned PDF OCR Fallback: Automatic optical character recognition via Tesseract.js in 5-page parallel chunks when no text layer exists.
  • 300+ Neural AI Voices: High-fidelity Microsoft Edge Neural TTS voices across 142 locales with customizable reading speed (0.5x to 2.0x) and inter-paragraph pause timing.
  • Smart Text Normalization: Intelligent cleanup algorithms that strip unwanted running headers, page numbers, URL clutter, and citations.
  • Batch Export: Synthesize entire chapters or full books into downloadable MP3 files or structured ZIP archives.
  • Privacy First: 100% local document parsing and IndexedDB storage—zero server uploads of personal files.

🛠️ How We Built It

Blob is built using modern web architecture and client-side processing pipelines:

  • Frontend & Framework: Next.js 16 (App Router) paired with React 19, Tailwind CSS v4, and Framer Motion.
  • Document Parsing Layer: pdfjs-dist for PDF layer extraction, mammoth for DOCX, jszip for EPUB/PPTX metadata, and tesseract.js for WebAssembly OCR.
  • Audio Processing Engine: A custom MPEG-2.5 frame synthesizer operating directly on raw array buffers to perform frame concatenation and insert silent audio buffers without re-encoding or introducing pops/clicks.
  • Storage & State: Web APIs with IndexedDB for local document storage and persistent offline application state.

🧮 Audio Frame Dynamics & Timing Model

To achieve natural reading cadences without audio distortion, Blob computes custom pause insertions dynamically using raw frame manipulation rather than time-stretching:

  • Total Audio Duration:
    T_total = Sum( t_synth(b_i) / rate ) + (N - 1) * Delta_t_pause

  • MPEG Frame Timing:
    Each MPEG frame contains S = 1152 audio samples. At a sample rate of f_s = 24,000 Hz, the frame duration is exactly:
    t_frame = 1152 / 24000 = 0.048 seconds (48 ms)

  • Silent Frame Calculation:
    The exact number of silent frames K prepended or appended between paragraph blocks for a delay tau is computed as:
    K = ceil( (tau * f_s) / S ) = ceil( tau / 0.048 )

By dynamically inserting K exact silent MPEG headers, Blob achieves seamless block transitions with zero re-encoding overhead.


⚡ Challenges We Ran Into

  1. Client-Side Memory Constraints During OCR: Running Tesseract WebAssembly workers across multi-hundred page scanned PDFs caused browser memory spikes and UI thread freezes. We solved this by designing a 5-page chunked pipeline with aggressive canvas garbage collection.
  2. Audio Artifacts in Frame Stitching: Naive concatenation of raw TTS audio streams caused distinct pop and click artifacts due to phase mismatch across MPEG frame boundaries. We implemented a frame boundary alignment algorithm that inspects MPEG frame sync bits (0xFFF) before inserting silent frame padding.
  3. Smart Header & Footer Striping: Document text layers often clutter reading flows with running page headers and footers. We developed a heuristic text normalizer that calculates text position variance and regular expression repetition across pages to strip non-narrative elements automatically.

🏆 Accomplishments That We're Proud Of

  • 100% Zero-Server Privacy: Successfully running complex multi-format document parsing, OCR, and audio synthesis completely client-side.
  • Massive Voice Selection: Unlocking access to 322+ global neural voices across 142 locales without requiring user API keys or subscriptions.
  • Lightning Fast Parsing: Sub-second parsing for standard 100+ page EPUB and text-layer PDFs using web workers.
  • Offline Storage: Full PWA compatibility with IndexedDB storage for seamless reading on mobile and desktop.

📚 What We Learned

  • WebAssembly Power & Limits: Deepened our understanding of WASM memory allocation when pairing PDF.js and Tesseract.js in modern Web Workers.
  • Audio Bitstream Manipulation: Gained hands-on experience inspecting raw MPEG frame headers and binary ArrayBuffers directly in JavaScript.
  • Accessibility & UX Design: Designing media controls and queue tracking that cater to visual impairment needs, focused on keyboard shortcuts and screen-reader accessibility.

🔮 What's Next for Blob

  • AI Summarization & Q&A: Integrating local LLMs (via WebGPU / Transformers.js) to offer instant chapter summaries and document Q&A before listening.
  • SSML Support: Allowing fine-grained control over voice emotion, pitch adjustments, and custom pronunciation dictionaries for complex technical terms.
  • Multi-Device Synchronization: Secure, encrypted peer-to-peer syncing of document reading positions across mobile and desktop devices.

Built With

Share this project:

Updates

Submission history