Inspiration

Most AI assistants are chat-first. You ask a question, get an answer, and still have to do the work yourself.

We wanted to change that.

What if you could simply tell an AI what you want, give it control of your desktop, and let it figure out how to complete the task?

Whether you're fixing a broken build, creating content, navigating applications, or completing a multi-step task, Cup Work can use the screen, keyboard, and mouse to work on your behalf.

Instead of: Open website → find the tool → copy information → switch applications → execute → verify

you can simply: Speak → Act → Verify

We wanted to move beyond the chatbot model and build an AI coworker that can understand your goal, operate your computer, adapt across applications, and work alongside you using general-purpose AI models.

That idea became Cup Work: an agentic AI desktop coworker built with Google ADK and Gemini.


What It Does

Cup Work follows a simple loop:

See → Explain → Act → Verify

It listens to your voice, understands your screen, explains what is happening, takes action on your Windows desktop, and verifies the result.

For example:

Broken container → identify the problem → explain it visually → request approval → execute the fix → verify the container is running.

Cup Work can:

  • 🎙️ Understand — Listen to voice directly through Gemini.
  • 👁️ See — Inspect applications using Windows UI Automation (UIA) and Gemini Vision.
  • 🖍️ Explain — Highlight problems and draw visual diagrams.
  • 🙋 Ask — Request Human-in-the-Loop (HITL) approval before critical actions.
  • ⚙️ Act — Interact with Windows applications using UI automation, mouse, and keyboard control.
  • ✅ Verify — Inspect the screen again instead of assuming the task succeeded.

It doesn't just tell you what to do. It shows you, asks you, does it, and checks the result.


How We Built It

Cup Work connects an Electron + React desktop client to a Python + Google ADK AI backend through local WebSockets.

Desktop Client

  • Voice Activity Detection (VAD) for natural voice interaction.
  • Streaming 24kHz PCM audio with real-time, expressive voice responses.
  • Transparent overlays for SVG whiteboards, screen highlights, and pointers.
  • Interactive task and command cards for communicating agent actions and results.
  • Cross-application screen highlighting and visual guidance.

AI Brain

  • A Root Agent routes tasks to the right specialist.
  • 7 specialized agents handle desktop automation, visual interaction, diagrams, research, clarification, command cards, and conversation.
  • Gemini 3.7 Flash via Vertex AI provides reasoning, vision, and tool execution.
  • Gemini 3.1 Flash TTS Preview provides real-time voice output.
  • SQLite stores preferences, active tasks, and time-aware memory.

Time-Aware Memory

  • Stores user preferences with timestamps.
  • Marks old preferences as expired instead of deleting them.
  • Helps the agent understand current and past preferences.
  • Provides better context during long-term interactions.

Controlled Agent Architecture

Each specialist agent receives only the tools required for its task. This reduces tool confusion, limits unnecessary actions, and makes the overall system easier to control and reason about.


Challenges & Engineering Solutions

  • UI Inconsistency: Not every application exposes reliable accessibility data. We use UIA when available and Gemini Vision as a fallback, using normalized coordinates.
  • Desktop Latency: A keyboard-first execution strategy and native shortcuts make actions faster and more reliable.
  • Voice Streaming: We built a chunk-coalescing buffer to prevent playback issues from small audio chunks.
  • Too Many Tools: Instead of giving one agent dozens of tools, Google ADK specialist agents isolate capabilities so each agent focuses on a specific task.

What Makes Cup Work Different

Cup Work combines:

Voice + Vision + Agents + Desktop Control + Human Approval + Time-Aware Memory + Verification

What we're most proud of is that these capabilities are not locked to a single application or website.

The same visual highlighter can work across a game, website, Google Cloud, or another desktop application. The same agent can explain something on screen, research information, and bring the result into Word, Notepad, or another application—including formatting, highlighting, and organizing the content.

Instead of building separate AI tools for every application, we built a centralized desktop AI layer that reuses the same capabilities across different environments.

One capability. Many applications. One AI coworker.


What's Next

  • macOS & Linux: Expand Cup Work beyond Windows.
  • AI Model Orchestration: Integrate specialized Google AI capabilities such as Gemini, Veo, and Lyria, so users can create, code, research, and generate media from one desktop workspace without switching between platforms.
  • Proactive Assistance: Detect build and terminal failures and suggest fixes automatically.
  • Community ADK Plugins: Let developers add specialist agents for tools like VS Code, Figma, and Blender.
  • Public Beta: Release simple installers and bring Cup Work to real users.

Our long-term vision is simple: tell Cup Work what you want to accomplish, and let it coordinate the right AI capabilities and desktop tools to get it done.

Built With

Share this project:

Updates

Submission history