Inspiration

Government benefit forms exist to help people, but the forms themselves are often the barrier. Dense language, unfamiliar fields, no help if English or French isn't your first language — the people who most need programs like Ottawa's Hand in Hand recreation and culture fee support are often the people least equipped to fight through the paperwork to get it: someone with low literacy, a newcomer still learning the language, an elderly resident unfamiliar with online forms, or just someone exhausted after a long shift who doesn't have the patience left to parse a government website.

We wanted to try removing the form entirely — not the program, not the eligibility criteria, just the interface. Governments already offer a form through multiple channels — paper, phone, in person, online — specifically for accessibility reasons. We wanted to prototype a fifth: the same form, the same legal process, the same backend, with a conversational layer in front of it that a government could accommodate alongside the channels it already has, not replace them with. Instead of reading a form and figuring out what to type where, you talk about your situation the way you'd explain it to a person behind a counter, and the structure gets built for you.

What it does

Open SpeakGov and, before you're asked to type or say anything, you can tap one button and hear what the form needs — in a real voice, in plain language, not a list of field labels you have to read first. Then you answer, out loud or by typing (same input, your choice, and voice never blocks the flow if it fails):

"Hi, my name's Zayd Houachmi, I live at 110 Dunbarton Court in Ottawa, I have two dependents, and my annual income is about 50 thousand dollars."

Gemini extracts the answers for a real program's eight-field form (modeled on the City of Ottawa's Hand in Hand recreation fee support) and they type themselves in live on screen. Leave out something required and it tells you, out loud if you were speaking: "Got it. I still need your address and your annual household income." Answer just that part and it merges in. You review what it heard — correcting a field is one tap — then confirm, and ElevenLabs reads the completed summary back to you before you're done. Once it's confirmed, you can print it or save it as a PDF to keep or bring to a service counter. Log in with Auth0 and your progress saves automatically, so a half-filled form is there when you come back.

Both directions are spoken, not just one. Reading text is never required to use the product start to finish — which matters, because a product that only reads your answer back to you still assumed you could read the question in the first place.

The whole thing works in English or French, chosen explicitly rather than guessed from your browser — Ottawa is officially bilingual, and a French-speaking user gets a fully French interface, not just a translated label here and there. And for screen-reader users, the page narrates what just happened ("Filled in: full name, address. I still need your income."), marks required fields, and respects reduced-motion settings.

How we built it

Next.js 16 (App Router, Turbopack) deployed on a Vultr VPS behind Caddy, which handles automatic HTTPS via Let's Encrypt for our GoDaddy domain, speakgov.com. The browser's Web Speech API does live speech recognition, with words appearing as you speak. Gemini 3.8 Flash does the structured extraction from raw speech/text, and on /any-form it also turns a pasted form into a field list. ElevenLabs voices every spoken moment: the "what do I need to say" prompt, the "I still need…" follow-up, and the readback. Auth0 gates a simple save/resume feature — one JSON blob per user, nothing more elaborate than the feature actually needs. PM2 keeps the Node process alive and restarts it on crash.

The build itself leaned heavily on Claude Code as a development partner — not just for writing code, but for the actual debugging work described below: reading raw API error payloads, diffing SDK source when documentation didn't match reality, and testing every fix against the live deployment rather than assuming it worked. Every commit in the repo reflects that collaboration honestly.

Challenges we ran into

Gemini silently dropping fields. Our first real extraction test came back with one field, wrong, out of four. Digging into the raw response showed the model was spending hundreds of tokens "thinking" about all four fields internally, then writing only one to the visible output before stopping. The fix was counterintuitive: mark every field as schema-required (with nullable types) instead of optional, which forces the model to actually consider each one instead of taking a shortcut. Verified with partial input, full input, and French input before trusting it.

Free-tier rate limits, twice. Testing burned through Gemini's 20-requests-a-day free tier not once but twice over the course of the build. The real fix was enabling billing (pennies per call at our volume); the code fix was adding a hard timeout, since a rate-limited request was taking 30+ seconds to fail on its own — long enough to look like a frozen page during a live demo instead of a clean error.

Speech recognition genuinely struggles with names. Live testing on an uncommon name ("Zayd Houachmi") produced a transcript where the recognizer misheard it as "Dave," the speaker tried to correct it by spelling it out loud, and because we'd set recognition to continuous: true to fix an earlier "cuts me off mid-sentence" bug, every retry got appended into one increasingly garbled blob. The fix wasn't more clever recognition — it was removing an auto-submit-on-stop behavior so the person can see what was actually heard and fix it before it goes anywhere, the same way a typo in the text box always could be.

Two language bugs from the same root cause. Both the speech recognizer and the text-to-speech readback were silently inheriting the browser's OS locale instead of respecting the explicit English/French toggle in the UI — so on a French-locale machine, the mic transcribed English speech as French, and separately, the "read it back" voice spoke French even in English mode. Same fix both times: never infer language, always pass the explicit user choice.

A Turbopack bug, confirmed against the framework's own source and issue tracker. Auth0's login route worked in every test except the deployed one, returning a 404. Chasing it down led through a confirmed, currently-open Next.js bug where Turbopack silently fails to populate the middleware manifest for the new proxy.ts convention specifically — the build claims success and even prints "Proxy (Middleware)" in its own summary, but nothing is actually wired up at runtime. The eventual fix mounts the same Auth0 logic as an explicit Route Handler instead of relying on Next's middleware auto-mounting, sidestepping the bug rather than waiting on it to be patched.

Accomplishments that we're proud of

A solo build that's fully deployed, on a real domain, with real HTTPS, and every integration genuinely working end to end rather than half-wired: Gemini, ElevenLabs, Auth0, Vultr, and a GoDaddy domain, in one coherent product rather than five bolted-together demos. Every bug listed above was found by testing against the live site with real input, not assumed fixed — including a maple-leaf mascot that took five failed hand-drawn attempts before we pulled the actual geometry from Canada's flag SVG and it worked on the first try.

What we learned

That the parts of a hackathon project people don't show off — a schema that forces a model to actually try, a rate-limit timeout, a language toggle that's honored everywhere instead of half the app — are usually where the real reliability comes from. And that when documentation and reality disagree, the fastest path is reading the actual source or the actual error, not guessing a second and third time.

We also learned something about voice products the hard way: late in the build we swapped the browser's recognizer for a server-side speech-to-text model. It was more accurate on names, but you had to stop talking before any words appeared. After trying it live, we reverted it within hours. For something you talk to, seeing your words appear as you speak mattered more than a slightly better transcript.

What's next for SpeakGov

Right now SpeakGov handles one form really well, on purpose — reliability mattered more than breadth for a weekend build. But to test whether the pattern generalizes, we added an experimental second mode at speakgov.com/any-form: paste the questions from any form, Gemini turns them into a field list (required marks, checkbox options and all), and the same voice flow fills it in. It works on simple forms today; multi-page and conditional forms are the gap. The natural next step is a front door that listens to someone's situation and routes them to the right form among several, which is a classification problem rather than an extraction one. We also built a WhatsApp channel through Twilio: the same form and the same "I still need your address" follow-up, just over chat, from any phone. It's tested on our server right up to the send step, but not live, because Twilio's trial won't forward incoming messages without a paid upgrade. Turning it on is a billing step, not an engineering one. Beyond that: translating extracted answers into an official language when someone speaks a language the form doesn't accept.

But the real next step isn't a feature — it's a conversation. We're not proposing that any government replace their forms or their systems with this. We're proposing the pattern: a conversational layer in front of an existing form, changing nothing about the form itself, that a government could accommodate as one more accessible channel alongside the ones it already offers. That's a much smaller ask than "adopt our app," and a much more honest one — and it's the actual goal behind building this at all.

Built With

Share this project:

Updates

Submission history