Inspiration

Opening YouTube costs four steps: reach for the mouse, find the window, click the address bar, type. The intent — "open YouTube" — takes half a second to say. The same is true of an idea you want to write down: you have to leave whatever you were doing and go find somewhere to put it. The interface is the tax, and you pay it a hundred times a day.

Every voice assistant we tried either lived inside one app, or sent everything we said to somebody else's server. We wanted one that controls the whole laptop and never leaves it.

What it does

You hold Ctrl+Alt+S anywhere — in a game, a PDF, your editor — say what you want, and let go. SaySo opens sites and searches, launches installed apps and folders, writes notes and reads them back out loud, sets reminders that keep ringing until you deal with them, and controls your browser tabs.

Press Esc while still holding the key and the recording is thrown away, because you usually know you said the wrong thing before you finish saying it. "Undo that" reverses the last change. And it learns your vocabulary: "when I say let's work, open notion and github" is one phrase and two windows from then on.

It is one system with three faces — a background desktop process, a live web console on localhost, and a Chrome extension. The console and the extension are views onto the daemon; close either one and every voice command still works.

How we built it

A global hotkey listener starts and stops a 16 kHz mono recording that goes straight into memory. faster-whisper transcribes it on the machine. A regex grammar maps the transcript to an intent in under a millisecond, and an executor carries it out. Everything is published on an event bus that the console and the extension both read over Server-Sent Events.

Three threads keep it responsive: the hotkey listener only starts and stops the recorder so it never blocks a keystroke, one worker does transcription and execution — which also serialises Whisper, since it is not safe to call twice at once — and Flask serves the console separately.

We picked the model by measuring rather than guessing. On the 4-core, 6 GB laptop this was built on: tiny.en took 1.7–2.2 seconds per command, base.en 3.4, and small.en 25–30 seconds, which would have destroyed a live demo. tiny.en ships.

Challenges we ran into

The first design was a single key, S for Speak. It would have swallowed the letter "s" in every text field on the machine.

Windows reports Ctrl+S as the control character \x13, not "s", so matching on characters failed silently. The listener matches virtual key codes instead.

Stopping the alarm killed the entire process — not an exception, the whole thing vanished and took the web server with it. The audio library drives one shared stream, and aborting it from a different thread than the one blocked on it takes PortAudio down at the C level. Every sound now writes to a stream it opened itself, in small chunks, and stopping only sets a flag the playing thread checks between them.

"Remind me in ten minutes" and "remind me to call mum" are the same three words. Rule order became load-bearing, and the tests now pin both readings so a later edit cannot quietly swap them.

Whisper never reports hearing nothing — it reports its best guess at what a cough might have been. That is how "Ah, luringy, luringy, luringy" ended up on screen as a failed command. Stutter loops and known hallucinations are now classified as noise and answered with "Did not catch that".

Accomplishments that we're proud of

It is honest about itself. The console labels every connector "works offline" or "needs internet" on the row you are about to switch on, rather than burying it in a README. The trade-offs are measured and written down, including the ones that went badly.

And the speech model is a component, not the product. Take it out and there is still a hotkey daemon, an audio pipeline, an intent grammar, an action executor, a note store, a scheduler and two live clients.

What we learned

That state a user can see should never be inferred from whether a thread is alive. That a help panel which runs real commands on click is a trap — ours set a timer nobody remembered asking for, and the alarm went off two minutes later. And that the interesting question about speech recognition is not whether it mishears you, but how quickly you can take it back.

What's next for SaySo

Streaming recognition, so the transcript appears while you are still speaking. More MCP connectors. And a Linux and macOS build of the daemon — the hotkey and audio layers are the only parts that are Windows-specific.

Built With

Share this project:

Updates

Submission history