Inspiration
Everyone in New York has filed a 311 complaint and watched it turn into a ticket number. Weeks later the status says closed. That word covers a repair, an inspector who knocked and got no answer, a visit where nothing was found, and a handoff to another agency. You never find out which one you got.
There's a field in the 311 dataset called resolution_description that says what actually happened, and almost nobody parses it. We classified it across 22 million records and found things like three out of four plumbing complaints to HPD closing with nothing found. We also thought about who needs to know that. It's the person with no heat at 11pm describing a cold radiator out loud, probably in a language the complaint form doesn't offer, so the front door of this thing is a microphone.
What it does
You hold the mic and say what's wrong. Your words show up on screen while you're still talking so you can fix anything that got heard wrong. Then 311derful tells you what will likely happen to that complaint based on real outcomes for your complaint type, your community board, and this time of year. The answer comes back written and spoken aloud in whatever language you used.
Behind one request: it maps your sentence onto NYC's official 311 taxonomy (a cold radiator is HEAT/HOT WATER → ENTIRE BUILDING under HPD), geocodes you to a community district, pulls the outcome distribution for that slice, and hands back advice plus a draft you can file yourself. Thin sample sizes get flagged LOW out loud. There's also an explorer view, for browsing citywide breakdowns without describing anything.
How we built it
The LLM never touches a number. It maps free speech onto the taxonomy, and it phrases a finished result in your language. Every statistic comes from an offline pipeline: pull the full dataset from Socrata month by month into DuckDB, classify every resolution description into an outcome class, roll it into a stats cube. At request time the forecast is a pure function over that cube, and the numbers arrive in the prompt already computed.
Voice runs through Vapi, used strictly as a transcriber. The second the call connects we send a control message muting the assistant, because otherwise it starts answering in its own words while we're computing the real answer. Vapi and the browser's Web Speech API return the same session object, so the mic component just picks a starter and the backend decides which engine is live. Web Speech can't detect spoken language, so that path shows a language picker and the Vapi path doesn't. The reply is spoken back through browser speech synthesis in the detected language, and typed input hits the exact same chain.
Challenges we ran into
Live transcription is messier than the docs suggest. A final transcript doesn't mean the speaker stopped, it means the transcriber hit an endpoint, so one sentence about a leaking ceiling shows up as "Okay." / "Um." / "There's water." Our first version overwrote the box each time and ate most of what people said. Ending a call also ejects you from the underlying Daily room and fires an error a moment later, which we had to learn to ignore so a normal hangup stops wiping a good transcript off the screen.
The classifier broke twice. Agencies rewrite their templates while you're not looking (NYPD changed its wording in November 2025, HPD twice before that), so we match short invariant fragments and report coverage per year instead of pooled. Worse, HPD's biggest template is "No violations were issued," which contains "violations were issued," and an early rule ordering silently flipped a million rows from NOTHING_FOUND to ACTION_TAKEN without throwing anything. There's a test pinning that order now.
Accomplishments that we're proud of
Classifier coverage stayed between 92.9% and 94.4% in every year from 2020 through 2026, checked year by year rather than averaged. Anything built on fewer than 30 records gets flagged LOW on screen and in the spoken reply, which most "AI plus city data" demos skip.
The voice path degrades all the way down: Vapi, then Web Speech with a picker, then plain typing if the browser has neither. And the hallucination question is boring here, because no model call in the system ever produces a statistic.
What we learned
"Closed" is doing a huge amount of quiet work in how open data gets reported, and none of that survives once you read the resolution text sitting in the next column. Splitting closure times by outcome changed our own read of the data more than once, since the fastest closures turn out to be the ones where nothing got fixed.
On the voice side, a transcript is a guess that keeps revising itself, and the UI has to show that honestly. Muting the assistant was the most important line we wrote, because the whole thing falls apart the moment a model improvises a number into someone's ear.
What's next for 311derful
We want the language handling to live entirely inside the voice pipeline so nobody picks a language before speaking. The backend already detects it from the transcript, so the picker only survives in the Web Speech fallback. We'd also like coverage across more agencies, plus alerting so a template change gets caught the week it happens.
NYC has no public write API for 311, so today we hand you a draft to submit yourself and say so on the page. The direction we're most curious about is flipping the tool toward community boards, where the question becomes which recurring complaints in a district are structurally unlikely to ever get fixed.
Built With
- ai
- claude
- duckdb
- fastapi
- google-visualization
- multilingual
- opendata
- vapi
- voice-to-text

Log in or sign up for Devpost to join the conversation.