💡 Inspiration
Access to information should be effortless, accessible, and private. Millions of students, visually impaired individuals, and multitaskers struggle to digest long-form text—whether it is scanned academic papers, PDF manuals, or eBooks. Existing text-to-speech tools often locked natural-sounding voices behind expensive subscriptions, failed when reading scanned images without text layers, or uploaded sensitive private documents to remote cloud servers.
We built Blob to eliminate these barriers: a 100% client-side, privacy-first document reader and audiobook creator that transforms complex documents into natural-sounding neural audio across 140+ languages—completely inside the user's browser.
🚀 What It Does
Blob converts text, PDFs, Word files, presentation decks, and scanned images into continuous high-fidelity audio:
- Multi-Format Ingestion: Instant parsing of
.pdf,.epub,.docx,.pptx, and.txtfiles. - Scanned PDF OCR Fallback: Automatic optical character recognition via Tesseract.js in 5-page parallel chunks when no text layer exists.
- 300+ Neural AI Voices: High-fidelity Microsoft Edge Neural TTS voices across 142 locales with customizable reading speed (0.5x to 2.0x) and inter-paragraph pause timing.
- Smart Text Normalization: Intelligent cleanup algorithms that strip unwanted running headers, page numbers, URL clutter, and citations.
- Batch Export: Synthesize entire chapters or full books into downloadable MP3 files or structured ZIP archives.
- Privacy First: 100% local document parsing and IndexedDB storage—zero server uploads of personal files.
🛠️ How We Built It
Blob is built using modern web architecture and client-side processing pipelines:
- Frontend & Framework: Next.js 16 (App Router) paired with React 19, Tailwind CSS v4, and Framer Motion.
- Document Parsing Layer:
pdfjs-distfor PDF layer extraction,mammothfor DOCX,jszipfor EPUB/PPTX metadata, andtesseract.jsfor WebAssembly OCR. - Audio Processing Engine: A custom MPEG-2.5 frame synthesizer operating directly on raw array buffers to perform frame concatenation and insert silent audio buffers without re-encoding or introducing pops/clicks.
- Storage & State: Web APIs with IndexedDB for local document storage and persistent offline application state.
🧮 Audio Frame Dynamics & Timing Model
To achieve natural reading cadences without audio distortion, Blob computes custom pause insertions dynamically using raw frame manipulation rather than time-stretching:
Total Audio Duration:
T_total = Sum( t_synth(b_i) / rate ) + (N - 1) * Delta_t_pauseMPEG Frame Timing:
Each MPEG frame containsS = 1152audio samples. At a sample rate off_s = 24,000 Hz, the frame duration is exactly:
t_frame = 1152 / 24000 = 0.048 seconds (48 ms)Silent Frame Calculation:
The exact number of silent framesKprepended or appended between paragraph blocks for a delaytauis computed as:
K = ceil( (tau * f_s) / S ) = ceil( tau / 0.048 )
By dynamically inserting K exact silent MPEG headers, Blob achieves seamless block transitions with zero re-encoding overhead.
⚡ Challenges We Ran Into
- Client-Side Memory Constraints During OCR: Running Tesseract WebAssembly workers across multi-hundred page scanned PDFs caused browser memory spikes and UI thread freezes. We solved this by designing a 5-page chunked pipeline with aggressive canvas garbage collection.
- Audio Artifacts in Frame Stitching: Naive concatenation of raw TTS audio streams caused distinct pop and click artifacts due to phase mismatch across MPEG frame boundaries. We implemented a frame boundary alignment algorithm that inspects MPEG frame sync bits (
0xFFF) before inserting silent frame padding. - Smart Header & Footer Striping: Document text layers often clutter reading flows with running page headers and footers. We developed a heuristic text normalizer that calculates text position variance and regular expression repetition across pages to strip non-narrative elements automatically.
🏆 Accomplishments That We're Proud Of
- 100% Zero-Server Privacy: Successfully running complex multi-format document parsing, OCR, and audio synthesis completely client-side.
- Massive Voice Selection: Unlocking access to 322+ global neural voices across 142 locales without requiring user API keys or subscriptions.
- Lightning Fast Parsing: Sub-second parsing for standard 100+ page EPUB and text-layer PDFs using web workers.
- Offline Storage: Full PWA compatibility with IndexedDB storage for seamless reading on mobile and desktop.
📚 What We Learned
- WebAssembly Power & Limits: Deepened our understanding of WASM memory allocation when pairing PDF.js and Tesseract.js in modern Web Workers.
- Audio Bitstream Manipulation: Gained hands-on experience inspecting raw MPEG frame headers and binary ArrayBuffers directly in JavaScript.
- Accessibility & UX Design: Designing media controls and queue tracking that cater to visual impairment needs, focused on keyboard shortcuts and screen-reader accessibility.
🔮 What's Next for Blob
- AI Summarization & Q&A: Integrating local LLMs (via WebGPU / Transformers.js) to offer instant chapter summaries and document Q&A before listening.
- SSML Support: Allowing fine-grained control over voice emotion, pitch adjustments, and custom pronunciation dictionaries for complex technical terms.
- Multi-Device Synchronization: Secure, encrypted peer-to-peer syncing of document reading positions across mobile and desktop devices.
Built With
- accessibility
- ai
- css3
- docx
- edge-tts
- epub
- framer-motion
- html5
- indexeddb
- javascript
- next.js
- node.js
- ocr
- open-source
- pdf.js
- privacy
- pwa
- react
- tailwind-css
- tesseract.js
- text-to-speech
- typescript
- web-audio-api
- webassembly
Log in or sign up for Devpost to join the conversation.