Why two patterns: the voice side needs LLM routing because natural speech doesn't fit a fixed graph. The monitoring side is the opposite — it must run the same way every time, so the order is enforced by the graph rather than by a prompt. You cannot classify before observing, or record before classifying. And because only the final stage has write access, a perception mistake still has to pass the grounding whitelist before it can reach the database.
Grounding: every tool call from every agent passes a before/after callback that validates arguments against whitelists — valid zones, valid violation types, valid severities. An agent physically cannot log a violation in a zone that doesn't exist.
Model separation: Gemini 3.5 Flash isn't a Live-class native-audio model, so the live voice model and the reasoning model are deliberately separate environment variables rather than one shared string. That separation is what let us upgrade the reasoning path to 3.5 without breaking voice.
Challenges we ran into
A one-line bug that muted the agent permanently. Voice worked at the start of every session and then went silent. Everything else kept working — transcripts, tool calls, routing — so there was no error to chase. The cause was a flag called audioSuppressed, set to true when Gemini fires an interrupted event (barge-in), and never set back to false anywhere in the codebase. The first time background noise triggered a false barge-in, AEGIS was muted for the rest of the session. There was even a _lastSuppressTime variable written but never read — evidence the original author intended a timed reset and never finished it. The fix was three independent reset paths: on turnComplete, on a 1.2-second timer as a fallback (the Live API sometimes drops turnComplete), and on reconnect.
An AudioContext leak that only appeared after a model change. Before finding the above, we chased a second real bug: the audio player created a brand new AudioContext on every WebSocket reconnect and never closed the old ones. Chrome caps concurrent contexts at about six, so the seventh reconnect threw an error that a catch block silently swallowed. It only surfaced when we tried a preview model that dropped sessions more frequently — more reconnects, faster exhaustion. Making the audio context a singleton fixed it. The lesson we actually took away: a bug that correlates with an unrelated change usually means the change altered frequency, not behaviour.
Making autonomy real, not theatre. Our first version had the vision agent "monitoring" — but it only ran while a WebSocket was open. That's not autonomy, that's a co-pilot. We rebuilt it as a FastAPI lifespan task with a separate Runner driving the ADK workflow pipeline via run_async() on a timer. Now closing the browser and running curl /api/monitor/status proves it: the counters advance with nothing connected.
Sessions that expired mid-demo. Live sessions kept closing cleanly every couple of minutes, and the client's fixed 1-second reconnect turned each drop into a storm of failed retries. Two fixes: enabling context_window_compression and session_resumption in the RunConfig so long voice sessions stop hitting the context limit, and replacing the fixed retry with exponential backoff that only resets once a session has actually proven healthy.
Resisting the demo lie. It was tempting to make the zone metadata look "real" — worker counts changing, names looking authentic. We pulled that back. The zone names and worker counts are clearly labeled as synthetic demo configuration; the vision detection, the pattern rules, and the risk scoring are what's real. Being upfront about that ended up making the actual capabilities more believable, not less.
What we learned
Deadlines force honesty. The thing we were proudest of a week in was the thing we cut on day four — a plan for on-device Gemma inference. Cutting it made room for the intelligence layer, which is what the project actually needed.
ADK is more expressive than one pattern. We'd only used LlmAgent with transfer_to_agent() before. Learning to compose with SequentialAgent and LoopAgent changed how we think about agent systems — some flows should be governed by the model, others should be governed by the graph. Picking the right one is a design decision, not a technical one.
Grounding is architecture, not prompting. Telling a model "don't hallucinate" doesn't work. What works is making hallucination structurally impossible — argument whitelists, deterministic scoring, stages with no write access. If the system can't produce the wrong output, you don't need to trust that it won't.
Diagnose by testing, not by reading. We spent time rewriting a ring buffer we were sure was broken. It wasn't — a simulation proved it handled 2.5 million frames without a fault. The real bug was six lines away in a flag that never reset. Reading code produces plausible theories; running it produces answers.
Real-time voice is unforgiving. There's no "let me check my notes." A tool call that takes 800ms feels like a hang. Every latency cost has to be paid down at build time, because there's nowhere to hide it at demo time.
What's next
- On-device Gemma for the confidence-gate stage, so obviously-clear frames never leave the edge
- Cloud Scheduler triggers for time-based autonomy — daily reports fire themselves at shift end
- Multi-site memory — patterns learned on one site inform risk baselines on another
- Real accuracy benchmarking against a labeled PPE dataset, so the review-accuracy percentage becomes a claim backed by ground truth
A last note
If AEGIS ends up on a real site someday, the moment that will matter isn't the demo — it's an officer sitting down after a break, asking "what happened while I was gone?", and getting a real answer.
Built With
- artifact-registry
- cloud-build
- cloud-run
- computer-vision
- css
- docker
- fastapi
- firestore
- gemini
- gemini-live-api
- google-adk
- google-cloud
- html
- javascript
- llmagent
- loopagent
- multi-agent
- opencv
- python
- sequentialagent
- uvicorn
- vertex-ai
- voice-ai
- websocket
Log in or sign up for Devpost to join the conversation.