Inspiration

The third time I tried to recreate my late mother's fried rice, I finally understood what I was actually losing.

It wasn't the recipe. The recipe exists. What doesn't exist is what she did before the wok ever touched heat, taking a small handful of dried chilies, crushing them slowly between her palms, rubbing them directly into the cold rice. The pressure. The duration. The exact moment she decided it was enough. Fifty years of pure instinct she never wrote down. I'll never taste it again.

Then I realized this isn't just personal. It's happening at a civilizational scale.

11,200 Americans turn 65 every single day through 2027. They retire, or they die, and the knowledge dies with them. Not the kind that lives in documents, the kind that lives in judgment. The NASA propulsion engineer who can hear an engine anomaly before any sensor fires. The semiconductor architect who knows which thermal edge cases will silently kill a chip two years post-launch because she's watched it happen twice. The reservoir engineer who knows exactly where the geological model is wrong because he's had a 30-year conversation with one specific basin.

McKinsey found knowledge workers already spend 20% of their workweek searching for information that was documented. The Panopto Workplace Knowledge and Productivity Report found that a 30,000-person company loses $72M annually from the productivity gap knowledge loss creates — and that 42% of institutional knowledge lives only inside individual employees and vanishes when they leave. The Center for American Progress puts the cost of replacing a highly trained expert at up to 213% of their annual salary.

Glean, Guru, Tettra, and the entire knowledge management category are built on one fatal assumption: that knowledge was written down somewhere. They find it faster. They surface it smarter. But not one of them can capture what the expert never knew to say. LLMs and RAG can't retrieve what was never written down. The instinct. The feel. The "I just know."

That's the gap the whole industry leaves open. That's why Servadium exists.


What it does

Servadium is a real-time multimodal knowledge transfer agent. AI that learns and teaches expert knowledge. It learns expert knowledge directly from experts, through live voice and video, and teaches it to learners in real time, with the same depth and visual fidelity the expert originally demonstrated.

It works in two modes:

Teach An expert performs a task while Servadium watches via camera and listens via microphone. The agent isn't a passive recorder, it's an active apprentice, asking clarifying questions, probing for edge cases, demanding the expert explain the things they usually skip because they've done it ten thousand times. When the session ends, Servadium semantically chunks the video, embeds everything multimodally, and distills the entire session into a permanent, structured skill.md knowledge artifact.

Learn A learner attempts the same task while talking to Servadium on camera. Servadium has the expert's full knowledge loaded as context, not just text, but the actual visual memory of how the expert moved, decided, and corrected. It coaches in real time via low-latency conversational audio. It sees what the learner is doing wrong before they do. And it renders AR overlays, arrows, circles, labels, directional markers, directly onto the learner's live camera feed the exact moment they deviate from the correct procedure.

The knowledge that lived only in one person's hands now lives in the system. And it can teach the next person, and the next, without the original expert ever being in the room.


How we built it

The core insight driving every architectural decision: this problem is fundamentally visual and conversational, not textual. The existing knowledge management industry failed because it tried to solve a visual, embodied problem with text retrieval. Servadium is built around that gap.

The Live Session Layer

Both Teach and Learn sessions run on Gemini Live (gemini-live-2.5-flash-native-audio) via the Vertex AI Live Connect API, streaming bidirectional PCM 16kHz audio and 1 FPS base64-encoded JPEG frames through a Python FastAPI backend via WebSockets. The backend sits in the middle to handle context injection, session state, and real-time recording to Google Cloud Storage.

In Teach mode, the system prompt turns Gemini into a relentless apprentice. In Learn mode, it becomes a real-time mentor with the expert's full knowledge loaded as context. When Gemini invokes the AR tool, a structured JSON payload, containing shape type (CIRCLE, ARROW, TEXT, FINGER) and exact screen coordinates, is forwarded to the frontend, where useAROverlay.ts renders it as a canvas overlay directly on the live camera feed.

Client-side, a ScriptProcessorNode splits the microphone signal: RMS volume drives UI animations sub-millisecond, while raw PCM is base64-encoded and streamed through WebSockets. To prevent AEC algorithms from muting the mic during Gemini's audio responses, the input processor routes through a GainNode set to zero before terminating at AudioContext.destination.

The Knowledge Modelling Pipeline

After a Teach session ends, the backend triggers an async Cloud Run Job via run_v2.JobsClient. The pipeline in backend/pipeline/agent.py runs on gemini-3.1-pro-preview and does five things:

  1. Analyses the session — reads the transcript, identifies what was clearly taught, what was ambiguous, and which visual moments carry information that words alone cannot
  2. Chunks the video semantically — generates a chunk plan with exact start_time_seconds and end_time_seconds per conceptual unit, then fires headless ffmpeg commands via subprocess to slice the raw recording
  3. Embeds everything multimodally — gemini-embedding-2-preview embeds raw video chunks, the transcript, and the skill.md together into a unified vector space in Vertex AI Vector Search; a video of someone demonstrating a hand motion and a text description of that motion become semantically adjacent
  4. Generates skill.md — distills prerequisites, hazards, structural notes, and the full timestamped transcript into a permanent structured knowledge artifact persisted in Cloud Storage
  5. Emits progress via Pub/Sub — every pipeline step publishes an event; backend/routers/webhooks.py catches the push and forwards it live to the user's waiting WebSocket connection

The Infrastructure

Everything runs on GCP:

  • Cloud Run — frontend (servadium-frontend) and backend (servadium-backend) containers, plus the ephemeral async pipeline worker job
  • Vertex AI Vector Search — production-grade semantic retrieval across all multimodal embeddings
  • Google Cloud Storage — raw .mp4 recordings via v4 signed URLs, semantic video chunks, and skill.md artifacts
  • Firestore — session state, user metadata, and knowledge chunk manifests
  • Cloud Pub/Sub — decouples the frontend from the background pipeline; bridges async completion events back into active WebSocket connections
  • Firebase Auth — user authentication and ephemeral token minting
  • Cloud Build — automated CI/CD via infra/cloudbuild.yaml

Challenges we ran into

Gemini Live connection stability The Live API is genuinely new and documentation is sparse. Getting a reliable bidirectional WebSocket proxy through a Python backend without dropped frames or hanging connections required significant debugging. The fix was setting {"ping_interval": None, "ping_timeout": None} in the google-genai SDK HTTP config to work around undocumented Cloud Run idle timeouts.

Video streaming saturation Pushing 24 FPS base64 video through the WebSocket architecture overwhelmed Gemini Live's input queues, causing stuttering and disconnects. Downgrading to 1 FPS via synchronous hidden canvas extraction preserved full spatial context while solving the throughput problem entirely.

Semantic video chunking Splitting video by conceptual meaning rather than timestamp is a genuinely hard problem. The pipeline has to reason about where one idea ends and another begins in a way no fixed algorithm can replicate.

Multimodal embedding pipeline Gemini Embedding 2 is brand new. Embedding raw video chunks at the right granularity, with the correct dimension size, and storing them correctly in Vertex AI Vector Search required working through undocumented edge cases.

Subprocess freezing during chunking Running headless ffmpeg inside the Cloud Run worker without stdout=subprocess.DEVNULL caused blocking process queues when pipe limits filled unexpectedly. Managing this with check=True and explicit error dumping solved the async segmentation failures.

Building alone under time pressure Every architectural decision had to be right the first time. The entire project was built from scratch in 24 hours.


Accomplishments that we're proud of

  • A fully working real-time multimodal pipeline: live video, live audio, live coaching — simultaneously, without any single layer becoming a bottleneck
  • AR overlays rendered live on a learner's camera feed, driven not by generic computer vision but by the specific expert's demonstrated knowledge
  • A knowledge system that embeds human expertise in a unified multimodal vector space — visual demonstrations and verbal descriptions retrievable by the same semantic query
  • A pipeline agent that validates its own chunking and synthesis before committing, catching gaps the original expert didn't know they left
  • Client-side RMS voice detection bringing UI responsiveness to sub-millisecond — without a single server round-trip
  • The whole thing rebuilt and shipped after losing five days of work, because the problem was too important to walk away from

What we learned

The hardest technical problem wasn't the AI. It was the architecture of real-time multimodal streaming, keeping video, audio, context injection, Pub/Sub events, and AR overlays all flowing simultaneously without any single layer becoming a bottleneck.

The deepest product insight was understanding why every existing solution failed: they all assumed knowledge was already externalized somewhere. The actual problem is capturing knowledge that has never been written down and was never intended to be. That requires a fundamentally different approach, not better search, but a system that can learn from a human the same way a human learns from another human. Through conversation. Through demonstration. Through being asked the right questions.


What's next for Servadium

  • Multi-session knowledge compounding — each new Teach session with the same expert deepens the artifact, catching contradictions, filling gaps, building a richer model over time
  • Cross-expert synthesis — merging knowledge artifacts from multiple experts in the same domain to surface consensus, disagreement, and edge cases only one of them has seen
  • Longitudinal learner tracking — the system knows not just what the expert knows but exactly where each individual learner is in their journey, adapting coaching to the specific gap
  • Enterprise knowledge graph — connecting all artifacts across a company into a living, queryable map of what the organization actually knows versus what it only thinks it documented
  • AR glasses integration — moving the coaching layer directly into the learner's field of view, not onto a separate screen

Built With

Share this project:

Updates

Submission history