Inspiration
I built Dev Caddie because technical knowledge is fragmented. We discover breakthroughs in Smart Feeds (RSS) but study them in deep-dive video lectures. These two worlds never talk to each other.
I wanted a Sovereign Agent that acts as a "Clubhouse Mentor"—an AI that doesn't just summarize, but walks the technical fairway with you. It stays with you all day: a morning audio briefing to start, a lecture companion while learning, and a feed assistant to discover. The goal was to eliminate the "latency gap" and "context drift," creating an agent that reasons across visual frames, transcripts, and external research with the immediacy of a human partner.
What it does (The Agentic Shift)
Sovereign Voice Agent (Gemini Live): This isn't a ‘chat-with-video’ wrapper. But a Sovereign Learning Agent that bridges your research and your study time into a single, persistent Knowledge Graph."
Lecture Caddie (Temporal Indexing): Deconstructs YouTube course content into a temporally-indexed knowledge graph. The agent knows exactly what slide you are looking at and can answer. Because these are pre-index the lecture, the agent isn't 'guessing' based on probability—it’s navigating a verified map of the content.
Smart Feed Orchestration: Ranks 100s of daily engineering posts by blending Gemini AI relevance with Community Signals (HN/Lobsters).
Deterministic Context Injection: Using InputTextRawFrame, the agent "Caddie-fies" the experience—dynamically injecting top-scored research articles into the live session based on the current lecture topic.
Live KPI Dashboard: Real-time observability tracking Agent health: TTFA, tokens, and grounding scores.
How we built it
We moved to a Sovereign Sidecar Architecture to separate UI state from real-time agentic execution:
The 4-Pass Pipeline (Ingestion): A deterministic Gemini 2.5 Flash pipeline that generates Markdown notes, visual slide markers, frame extractions via ffmpeg, and a final temporally-mapped concept graph.
Pipecat + Gemini Live: A GCE VM handles persistent WebRTC connections via Daily.co. We implemented Hard Anchoring where the YouTube currentTime is synced to the agent's system prompt to prevent context drift.
Agentic Grounding: We use LLMRunFrame Nudges and InputTextRawFrame to force-inject external research into the agent’s active memory, ensuring the agent remains proactive.
BigQuery & Redis: BigQuery stores the massive scored corpus; Redis handles the "Continuity Packet" for 24h conversational continuity.
Challenges we ran into
The Dual-Scoring Problem: Single-dimension AI scoring treats obscure blog posts the same as viral engineering breakthroughs. By combining Gemini's personal relevance scoring with Community Signaling (HN/Lobsters), the agent delivers field-tested recommendations.
VAD-Offset Drift: In long technical sessions, Voice Activity Detection often desyncs. We mitigated this by implementing manual context refreshes via LLMRunFrame to keep the agent responsive and sovereign.
Temporal Precision: Mapping a Knowledge Graph to a 2-hour video without re-sending tokens was hard. We moved to a "Marker-Pass" system where Gemini inserts metadata without re-processing raw video bytes.
Accomplishments that we're proud of
Deterministic Injection: Successfully blending a curated research feed into a live voice session without breaking the "Caddie" persona.
Dual-Scoring Moat: A robust filtering system that resists "obscure post" bias by requiring social proof before AI processing.
Cost-Optimization: Article deduplication by URL hash ensures Gemini only scores new articles each day — already-scored content is skipped entirely. Community signals (HN/Lobsters) act as a natural quality filter before AI scoring, and decay-based rescoring refreshes trending signals without re-invoking the AI.
Solo Build & Production Ready: A complete end-to-end agentic architecture—from RSS ingestion to WebRTC voice—built entirely by one person.
What's next for the Agent
Topological Memory: Expanding the knowledge graph into a 3D memory map where the user can "fly" through related technical concepts.
Agentic Agitation Scoring: Developing a "Concept Density" score to trigger proactive agent interruptions during complex lecture segments. (Roadmap)
Multi-Agent Personas: Switching "Caddie" archetypes (e.g., The Skeptical Reviewer vs. The Supportive Tutor) based on user goals.
Log in or sign up for Devpost to join the conversation.