Inspiration

I've used a lot of AI agents. Every single one makes the human do the work.

You craft the prompt carefully. You wait. It misunderstands. You rephrase. You wait again. That's not autonomy — that's a smarter search bar.

I wanted to build something that flips the contract entirely. Where the human speaks once, naturally, and the agent figures out the rest. Where you can interrupt mid-sentence, change your mind, go off-topic — and the agent keeps up without missing a beat.

That's what inspired Rio. Not a feature. A different philosophy of what human-agent interaction should feel like.


What it does

Rio is a fully autonomous voice agent that takes one instruction and completes it end-to-end — across applications, across modalities, without follow-up prompts.

It listens like a human. Rio uses Gemini Live API for native end-to-end audio — not a STT → LLM → TTS pipeline. Real voice in, real voice out, with true interruption handling via Silero VAD. When you start talking mid-response, Rio stops instantly and listens.

It sees your screen. Pillow and RapidOCR capture live screen frames and stream them to the cloud. Rio knows what's open, what's on screen, and can act on it.

It acts across your applications. Using the Google Workspace CLI and ADK tool orchestration, Rio chains actions across Gmail, Calendar, Sheets, and Docs — from a single spoken instruction. Say "log this complaint as a support ticket" and a row appears in Google Sheets in real time, with ticket ID, severity, category, and timestamp — extracted from your voice.

It understands when you're struggling. A scikit-learn classification model runs in real time, predicting from user behavior patterns — errors, retries, hesitation signals — when a user is stuck. Rio adapts its response style proactively, before you even ask for help.

It remembers. ChromaDB with Gemini Embedding 2 gives Rio semantic memory across sessions. It doesn't start from zero every conversation.


How I built it

Rio runs as a split system — local machine and cloud — connected by a binary WebSocket protocol I designed specifically for low-latency multimodal streaming:

$$\text{frame} = \underbrace{0\text{x}01}{\text{audio}} \cdot \text{PCM}{16}^{16\text{kHz}} \quad \bigg| \quad \underbrace{0\text{x}02}_{\text{vision}} \cdot \text{JPEG}$$

Local machine owns all hardware: mic capture, screen frames, UI actions, file system, browser automation via Playwright.

Cloud Run owns all thinking: Gemini Live session, ADK tool orchestration, model routing, and the ToolBridge pattern that dispatches tool calls from the model back to the local executor over WebSocket.

Model stack:

  • gemini-live-2.5-flash-native-audio — real-time voice I/O
  • gemini-2.5-flash — text reasoning and tool calls
  • gemini-2.5-pro — escalation for complex multi-step tasks
  • gemini-embedding-2-preview — semantic memory
  • scikit-learn classifier — real-time struggle detection from user behavior patterns

Rate limiting uses a token bucket with 4 degradation levels:

$$\text{RPM}_{\text{effective}} = \min\left(30,\ \frac{\text{tokens remaining}}{\Delta t}\right)$$

When the bucket drains, Rio degrades gracefully — reducing tool call frequency before ever dropping the voice connection.

Google ADK powers the tool orchestration. Every tool is closured per connection (the ToolBridge pattern), so each user session gets isolated, stateful tool execution without shared state bugs.


Challenges I ran into

The native audio routing problem. Gemini's native audio model (gemini-live-2.5-flash-native-audio) is unreliable for function calling. I had to route all tool execution through a text-based orchestrator via live_model_tools=False — which means there's a hidden probabilistic parsing layer between the model's intent and the tool dispatch. Getting that boundary clean without latency spikes took significant iteration.

Binary protocol design. Multiplexing audio and video over a single WebSocket without head-of-line blocking required a priority queue architecture — audio frames at priority 0, JPEG frames at priority 1. Without this, a large screen capture would delay audio by hundreds of milliseconds.

Interruption handling under load. Silero VAD running in a sounddevice callback thread needs to hand off into the asyncio event loop without blocking. Getting call_soon_threadsafe wired correctly so VAD-triggered interruptions felt instantaneous — not 200ms delayed — was harder than expected.

Session resumption. Cloud Run can cold-start mid-conversation. InMemorySessionService means a restart drops the Gemini Live session handle. I had to design reconnect logic on both the local and cloud side so the conversation context survives a server restart transparently.

Struggle detection integration. Training the scikit-learn model on meaningful behavioral signals — distinguishing genuine struggle from normal interaction variance — required careful feature engineering. Hesitation alone isn't struggle. The combination of retry frequency, error rate, and response latency patterns together form a reliable signal.


Accomplishments that I'm proud of

  • Built a genuinely interruptible voice agent — not simulated, not scripted. Real VAD, real native audio, real sub-100ms interruption response.

  • Designed a binary WebSocket protocol from scratch that multiplexes audio, video, and tool RPC over a single connection without blocking any lane.

  • The ToolBridge pattern — decoupling model-side tool calls from local execution entirely, with per-call timeout policies and result validation. Clean enough that adding a new tool is 10 lines of code.

  • A scikit-learn struggle detection model running live, classifying user behavior patterns in real time to trigger adaptive agent responses — before the user asks for help.

  • Cross-application autonomy over Google Workspace from a single voice instruction, with real visible output judges can verify in real time.

  • A 4-level graceful degradation model that keeps the agent functional under API rate pressure — something most hackathon projects ignore entirely.


What I learned

Architecture beats features every time. The split local/cloud design — local machine for sensing and acting, cloud for thinking and coordinating — was the hardest decision and the best one. It's what makes Rio's latency acceptable and its actions genuinely powerful.

Native audio is a different paradigm. STT → LLM → TTS pipelines feel like AI. Native audio feels like a conversation. The difference isn't technical — it's perceptual. Users respond completely differently.

Graceful degradation is not optional. Free-tier Gemini API limits are real. Building degradation in from day one, rather than as a patch, is what separates a demo that survives a live presentation from one that crashes at the worst moment.

The ToolBridge pattern generalizes. Any architecture where a cloud model needs to execute local actions benefits from this pattern. It's not Rio-specific — it's a reusable primitive I'll carry forward.

Behavioral ML belongs inside the agent loop. Running a prediction model alongside an agent is useful. Wiring it into the agent's decision loop is transformative. The gap between those two is the next frontier.


What's next for Rio Agent

Rio v1 proves the architecture. What comes next is the full vision.

The current build is a single powerful agent. The next version is a multi-agent system — four specialized sub-agents coordinated by a central orchestrator:

  • Orchestrator — ADK + Gemini 2.5 Pro, the brain that plans and delegates
  • Live Agent — Gemini Live API, owns all voice and vision interaction
  • UI Navigator — Playwright + DOM control + Computer Use API, replaces pyautogui with true visual grounding on any screen
  • Creative Agent — Gemini interleaved output + Imagen 3, generates mixed-media responses in a single fluid stream

These communicate via A2A protocol with Agent Cards — making Rio discoverable and composable with other agents in Google's ecosystem, not just a standalone tool.

Specific technical milestones:

Computer Use API — replacing pyautogui with gemini-2.5-computer-use-preview for true visual UI grounding. Click any element on any screen without pre-mapped coordinates.

Struggle detector as ADK streaming tool — the scikit-learn model is already live and predicting user struggle in real time. The next step is rewiring it from a standalone pipeline into a proper ADK streaming tool connected directly to LiveRequestQueue — so Rio's behavioral adaptation happens inside the agent loop, not outside it.

Production memory — replacing InMemorySessionService with persistent session handles and Vertex AI Vector Search for multi-user scale. Rio currently remembers within a session. The goal is remembering across weeks.

The real goal: Rio shouldn't be an app you open. It should be infrastructure that runs in the background — voice-activated, screen-aware, acting across your digital environment whenever you need it.

One human. One instruction. Rio does the rest.

Built With

Share this project:

Updates

Submission history