Inspiration
I noticed most voice-note apps quietly depend on a cloud server to do the actual transcription — which means no signal, no notes, and your voice recording is sitting on someone else's servers the whole time. That bothered me more than it probably should have. I wanted to see if I could build something where the "smart" part happens entirely on your own device, so it works whether you're in a dead zone or just don't want to hand your voice over to a company every time you jot something down.
What it does
EchoNotes lets you record your voice and get an instant, accurate transcript — entirely inside your browser tab. You tap record, speak, tap stop, and your words appear as a saved note in a running "notes reel," complete with an automatically generated title and timestamp. You can copy or delete any note. The core trick is that once the speech model has loaded once, the entire transcription pipeline runs locally: no audio, no transcript, and no note ever leaves your device, and it keeps working even with your internet turned off completely.
How we built it
I used vanilla HTML, CSS, and JavaScript with no build step and no backend at all. Speech-to-text is powered by Whisper-tiny.en, running through Transformers.js, which executes the model directly in-browser via WebGPU (with a WASM fallback). The browser's native MediaRecorder and AudioContext APIs handle capturing the microphone input and resampling it to the 16kHz mono format Whisper expects. Notes are saved with localStorage, so everything — audio capture, AI inference, and storage — happens client-side. I designed the interface around a warm cassette-deck/VU-meter aesthetic, with a live animated waveform that reacts to your voice while recording, to make the "you're talking to a local device, not the cloud" feeling tangible rather than abstract.
Challenges we ran into
The biggest early hurdle was getting audio into the right format for Whisper — raw microphone input from MediaRecorder isn't natively 16kHz mono, so I had to build a resampling step using an OfflineAudioContext to convert recorded audio into the exact Float32 array shape the model expects. I also initially passed a language parameter meant for multilingual Whisper models into the English-only whisper-tiny.en checkpoint, which caused transcription to fail outright — a good reminder to match model configuration precisely to the specific checkpoint you're using, not just the general model family. Getting the VU-meter visualization to react smoothly to live audio in real time, without janky animation, also took some tuning of the analyser node's FFT settings.
Accomplishments that we're proud of
I'm proud that this is a genuinely complete, working product rather than a proof-of-concept — you can turn off your wifi entirely and it still works, which is the whole point proven live rather than just claimed. I also like that the entire experience needed zero backend infrastructure: no server, no API keys, no hosting costs, just a static site anyone can clone and run. And I pushed to make the interface feel intentional and distinct rather than a bare-bones demo, since I think tools that protect people's privacy should feel trustworthy and polished, not like an afterthought.
What we learned
This was my first real hands-on project with on-device/local AI inference rather than calling a cloud API, and it changed how I think about the trade-offs in AI product design — running a model client-side means thinking hard about model size, load time, and format (ONNX specifically) in a way that calling a hosted API never forces you to. I also got much more comfortable with lower-level browser audio APIs (MediaRecorder, AudioContext, OfflineAudioContext) that I'd never needed before as a front-end developer.
What's next for EchoNotes
I'd like to add a small on-device summarization step so long recordings get a short auto-summary alongside the raw transcript, explore a multilingual Whisper variant for non-English speakers, and add basic search across saved notes. Longer-term, chunked transcription for multi-minute recordings would make it practical for things like full lecture capture, not just short voice notes.
Built With
- css
- html
- javascript
- localstorage
- mediarecorder-api
- onnx
- transformers.js
- web-audio-api
- webassembly
- webgpu
- whisper
Log in or sign up for Devpost to join the conversation.