Inspiration

Most people know what 911 is. Far fewer have heard of 311.

311 is to city problems what 911 is to emergencies: the number you call about the pothole, the flooded underpass, the traffic light stuck on flashing red, or the tree limb resting on a power line.

The people who answer are skilled professionals. A 311 dispatcher decides whether that limb outranks the flooding, which crew can actually do the job, and when a road has to close. That kind of judgment takes years to build. Yet a handful of dispatchers often covers a whole city, and most of their shift goes to work that doesn't need them. They read the fifth call about the same pothole. They decide, for the hundredth time this week, that a water main break goes to the water crew. They lose the dangerous report somewhere in a pile of faded crosswalks. And they rarely have time to tell the resident anything, so the person who called assumes the city ignored them.

We didn't want to replace the dispatcher. We wanted to take that pile off their desk, so their expertise goes to the calls that genuinely need a person. That's why we built 311.ai for the Professional Agents track.

What it does

311.ai is a team of AWS Strands agents that works underneath a city's 311 desk. A resident reports a problem by voice, and within one flow it becomes a geolocated incident with a work order, a recommended crew and a spoken confirmation. The agents handle routine dispatch on their own and bring a person in only when a real decision has to be made.

  • Live voice intake (Amazon Nova Sonic + Amazon Polly). A resident speaks and Amazon Nova Sonic on Bedrock transcribes them live. Nova Lite pulls out the category, severity and address, and Amazon Polly speaks the confirmation back. One intake function serves the live mic, posted transcripts and the MCP surface, so a report gets the same handling whichever way it comes in.
  • A coordinator with specialists (Strands Agents SDK). A Dispatch Coordinator on Amazon Nova Pro holds three specialists as callable tools: Public Works, Sanitation and a Resident Liaison. It can also decline to involve any of them and answer directly, which is what happens for most traffic. Only cross-department work fans out, like a water main break whose access is blocked by dumped debris.
  • 19 tools that change real state. The agents search, dedupe, score crews, open, amend and cancel work orders, assign crews, post field updates and draft resident messages. Each of those actions writes to Redis, and the board reflects it on the next poll. Nothing is canned.
  • Semantic deduplication (Amazon Titan + Redis vector search). Every report is embedded with Amazon Titan Text Embeddings and compared by meaning against open incidents, then filtered by distance. "Water gushing into the street" matches a report filed as "line fracture", so five calls about one problem become one job.
  • Auditable crew matching and risk ranking. Crews are scored on five weighted factors: specialisation (40), true distance (25), track record (20), workload (10) and readiness (5). The full breakdown goes into the work order's reasoning trace. Escalation risk is a deterministic score built from severity, age, category volatility and hotspot density. The LLM writes the explanation, but it never sets the order.
  • A clear autonomy boundary. Jobs below Critical severity, with a free crew and inside the cost ceiling, go out automatically. Critical jobs, busy crews and out-of-budget spend are held, and the card says why. When an agent reaches a call it shouldn't make alone, request_human_decision puts the question, the options and its recommendation in a supervisor queue, and the agent stops. For us, success is a board that keeps moving while that queue stays empty.
  • Work a small desk can't do. A duty agent sweeps the queue on a timer, including overnight, and triages the riskiest unassigned incidents first. A follow-up pass chases approved jobs that stopped moving, with patience that depends on severity: 4 hours for Critical, 12 for High, 48 for Medium and 96 for Low.
  • Live fleet dispatch (Amazon Location Service). Approving a job sends the right vehicles, like bucket trucks, patch trucks or sewer cleaners, along real truck routes with live traffic. They stop at real traffic signals from OpenStreetMap. The fleet is seeded and its movement is simulated and time-compressed, and the map says so.
  • Field Review. Crews post to the field log, upload photos with SHA-256 checksums stored in Amazon S3, and record on-site audio that Nova Sonic transcribes onto the work order.
  • Labelled incident illustrations (Stability AI on Bedrock). A generated image shows the kind of situation reported, so a dispatcher can judge severity quickly. The label "not the actual site" is printed inside the frame, so it survives a cropped screenshot.

How we built it

  • Agents: AWS Strands Agents SDK using the agents-as-tools pattern, on Amazon Bedrock (Nova Pro for reasoning, Nova Lite for fast extraction). A dispatch turn is capped at 180 seconds and 12 loop iterations. TokenRouter is only a fallback when a Bedrock turn fails, and the agent code is the same on both paths.
  • Amazon Bedrock AgentCore: The same agent team is deployed on AgentCore Runtime. AgentCore Memory keeps a conversation per incident, so follow-up directives keep their context. AgentCore Observability sends every agent turn, tool call and Bedrock request to CloudWatch GenAI Observability as a trace.
  • Voice: A 16 kHz mono PCM16 browser capture streams over a WebSocket to Amazon Nova Sonic, with Amazon Transcribe as an option and Amazon Polly for speech out.
  • Backend: FastAPI with async REST and WebSocket endpoints, a background sweep loop, and an MCP server so Claude Desktop or Cursor can drive the same dispatch tools.
  • Data: Redis Stack is the source of truth. RedisJSON stores incidents, work orders and crews, RediSearch handles vector KNN and geo filtering, and RedisTimeSeries holds operational metrics. MongoDB Atlas is an optional mirror, and Amazon S3 stores evidence.
  • Frontend: React 19, TypeScript, Vite and Tailwind. The map uses MapLibre GL with OpenFreeMap tiles, so it needs no API key, and incidents are drawn as a GPU-rendered GeoJSON layer.

Challenges we ran into

  • Voice that "worked" but heard nothing. Our first capture sent WebM/Opus from MediaRecorder. The speech engines accepted it without error and transcribed pure silence. The button worked and the call completed, but nothing was ever filed. Switching to raw 16 kHz PCM16 fixed it, and the same capture now feeds both Nova Sonic and Transcribe.
  • AWS errors that pointed the wrong way. Transcribe returned SubscriptionRequiredException, which looks like a credentials problem but is really an account state you can't change from the CLI. We moved live speech-to-text to Nova Sonic, which uses the Bedrock access we already had.
  • A dependency that broke only in the cloud. strands-agents-tools ships a top-level module called strands_tools, the same name as ours. Locally our file won. On AgentCore Runtime the package won, and every tool call failed. Separately, amazon-transcribe pins an awscrt version that conflicts with botocore[crt], which made a clean build impossible. We untangled both.
  • Search that went quietly wrong. At one point, adding AWS credentials changed the embedding model to one with the wrong vector width. Search kept running, but every incident became equally distant from every other. Now embeddings resolve separately from chat and are checked against the index dimension.
  • Three intake paths that drifted apart. Early on, the live mic, transcript upload and phone agent each had their own copy of intake, with different ID formats and only one assigning a crew. We merged them into a single ingest_voice_report.
  • Agents that deferred to the person who had just decided. When an operator typed "assign a crew", the agents would escalate that same decision back to them. Operator directives are now marked as authority already granted. The narrow boundary still applies: closures, pulling crews off higher-severity work, hazardous disposal and out-of-budget spend.

Accomplishments that we're proud of

  • The agents stop when they should. During development, agents used request_human_decision without being prompted to. One refused to report a crew as en route because it couldn't verify the work order behind that status.
  • A follow-up that diagnosed the problem. When the follow-up pass found a Critical flooding job that had sat idle for 58 hours, the agent first checked whether the crew was tied up on something worse. They weren't, so it posted a chase that named a backup crew five minutes away and left the stall clock running in case nobody answered. On the next job, it noticed one inspector held three stalled jobs and identified overload rather than a priority conflict.
  • Every step is auditable. Crew choices carry their score breakdown, risk ranking is reproducible, and each work order keeps a reasoning trace.
  • Honest about what's simulated. Simulated fleet movement, generated illustrations and seeded data are all labelled as what they are.
  • Built fully on AWS. The system runs on Strands, Bedrock, AgentCore Runtime, Memory and Observability, Nova Sonic, Polly, Titan, Location Service and S3.

What we learned

  • Autonomy needs a boundary. Giving agents a tool to escalate and stop made them more useful, because people can trust what they do on their own.
  • Let the LLM explain the order, not set it. A priority queue that reshuffles between identical runs is something an operations team will stop trusting. Deterministic scores with LLM-written explanations gave us both consistency and readable reasoning.
  • Agents-as-tools fit quiet systems. A graph runs every node, and a swarm gives up control. A coordinator that can decide not to involve anyone keeps routine work to a single hop.
  • Silent failures are the worst kind. Silent audio, equidistant vectors and a shadowed module all looked fine on the surface. Our preflight scripts now report each AWS service on its own line with the reason and the fix.

What's next for 311.ai

  • A real city system of record. We want to connect intake to a city's actual 311 platform through Open311 instead of seeded data.
  • Real fleet telemetry. We want to replace simulated vehicle movement with live GPS from city fleet vehicles.
  • A phone line. We want residents to be able to call a real number through Amazon Connect and reach the same Nova Sonic intake.
  • A mobile app for crews. We want a phone-first Field Review for crews, with offline capture and voice notes.
  • Predictive maintenance. We want to use recurring hotspots and incident history to flag infrastructure likely to fail before anyone calls.

Built With

Share this project:

Updates

Submission history