Live Shopping Agent — Project Story
Inspiration
Live shopping is a multi-hundred-billion-dollar industry, but it has a fundamental scaling problem: it requires human hosts to be present, energetic, and on-camera around the clock. Small and medium-sized brands simply cannot afford to run 24/7 live streams, even though the data shows that engagement and conversion rates are dramatically higher during live video compared to static product pages.
We were inspired by a simple question: what if a brand could have an always-on, AI-powered shopping host that never sleeps, never gets tired, and can engage with every single viewer by name? The explosion of real-time AI voice generation (Gemini Live API) and fast multimodal models made this feel suddenly achievable — not as a novelty demo, but as a genuinely production-viable system.
We also drew inspiration from the architecture of professional broadcast systems, where switching between content sources must be seamless and interruption-free. The challenge of applying those same reliability standards to a fully autonomous AI system is what made this project technically exciting.
What it does
Live Shopping Agent is an autonomous AI live-stream shopping host that runs 24/7 on platforms like Twitch without any human involvement.
At its core, the system operates in two modes that it switches between automatically:
- Ad Loop Mode — When no viewers are present, the stream quietly plays a rotating playlist of product ads. Zero compute is wasted on AI inference.
- AI Live Mode — The moment a viewer joins or sends a message, an AI host comes alive. It greets viewers by name, pitches products, answers questions, and keeps the energy up — all with a real-time generated voice streamed directly to the broadcast.
The AI host is actually two cooperating AI systems:
AI Director (Gemini 2.5 Flash) — A fast, cost-efficient model that monitors the chat every 5 seconds, triages incoming viewer events, and decides what the host should say next. It handles context like which products have already been pitched, which viewers are most engaged, and when it's time to run autonomous filler content if chat goes quiet.
AI Host (Gemini Live API) — The face (and voice) of the stream. It receives a
DirectorInstructionand opens a persistent Gemini Live session to produce a flowing, real-time spoken response — audio streamed continuously to the broadcast in precisely timed 20ms chunks. While speaking, it responds in the live chat simultaneously, tagging the viewer by name so the reply reaches both audio listeners and anyone reading chat. What makes it genuinely powerful is its tool-calling capability: mid-sentence, without breaking flow, it can call any of these live tools:get_current_product/get_next_product— know what's being showcased and advance the lineupget_product_details— pull full product info (name, description, attributes) by ID or SKUget_product_price— fetch exact price, discount percentage, and computed final priceget_payment_link— retrieve the purchase URL and post it directly in chatget_viewer_info— look up a viewer's join time, message history, and interaction count for personalized greetingsget_viewer_count— reference the live audience size naturally in conversationget_ad_playlist— know what ads are in rotationsend_chat_message— post text responses, payment links, or callouts to chat in real timeshow_product_on_screen— trigger the streaming engine to swap the visual feed to a product's promo video or imageclear_screen— return the visual feed to the default background
The result: when a viewer asks "how much is it?", the host answers aloud with the exact price and discount, then posts the payment link in chat — all without any scripting or human involvement.
The entire media pipeline — video encoding, audio mixing, and RTMP push to Twitch — is handled by a dual-process FFmpeg architecture that allows content to swap seamlessly without ever interrupting the stream.
How we built it
The system is structured as five distinct, single-responsibility layers:
| Layer | Responsibility |
|---|---|
| Channel Gateway | Abstracts Twitch IRC / OAuth2; emits platform-agnostic PlatformEvent objects |
| Broadcast Orchestrator | Traffic controller; manages mode switching and routes events + media bytes |
| Streaming Engine | Dual-process FFmpeg; encodes and pushes H.264/AAC to RTMP endpoint |
| Content Execution | Ad unit and AI Live unit; isolated content producers |
| Session Manager | PostgreSQL-backed source of truth for all shared state |
Key technical decisions:
Dual-process FFmpeg architecture — The single hardest problem was eliminating stream interruptions during mode switches. Starting a new FFmpeg process causes a 3-5 second black frame gap that Twitch penalizes. We solved this by splitting FFmpeg into a persistent output process (runs the entire session) and a swappable source process (replaced on each transition), joined by a Python relay thread that keeps the named pipe alive during swaps.
Tick-based Director — Rather than reacting to every single viewer event, the Director batches all events over a 5-second window and sends them together to Gemini Flash. This reduces API calls by an order of magnitude, gives the model better context, and produces more coherent instructions.
Gemini Live tool calling — The AI Host uses Gemini Live's function-calling capability to act on live data mid-response. It has a full tool suite: fetch product details, get the current or next product, look up precise pricing and discounts, retrieve payment links, query viewer history, check live audience count, read the ad playlist, post to chat, swap the on-screen visual to a product image or promo video, and clear the screen. When a viewer asks "how much is it?", the host calls
get_product_price()and weaves the exact figure into speech; when they're ready to buy, it callsget_payment_link()and posts the URL directly in chat.Pre-warming — When transitioning from ad mode to AI live mode, the AI components begin initializing in the background while ads continue playing. The switch only completes once the AI Host has produced its first audio chunk, guaranteeing a zero-interruption handoff.
Rate-controlled audio — Audio is delivered in precisely timed 20ms chunks at 24kHz mono s16le to prevent buffer build-up and audio/video desync over long sessions.
Stack: Python 3.11, FastAPI, PostgreSQL + async SQLAlchemy + Alembic, Gemini 2.5 Flash, Gemini Live API, FFmpeg, Twitch IRC, Docker + Docker Compose.
Challenges we ran into
Stream interruption on mode switches was the most painful problem. Every time we killed and restarted FFmpeg to switch from ads to the AI host, Twitch received a gap in the stream. It took significant experimentation to arrive at the dual-process architecture with a persistent output process and a Python-managed named pipe relay. The details are documented in Stream_Interruption_Fix.md.
Audio/video sync over long sessions was subtle but damaging. Small timing drift in audio chunk delivery would compound over hours, eventually causing the host's voice to lag noticeably behind the video. Rate-controlling audio output to exact 20ms intervals at 24kHz resolved this.
Coordinating two AI systems without creating feedback loops or race conditions required careful design. The Director and Host share a single instruction queue with strict producer/consumer discipline. The Host never writes back to the Director mid-cycle, and the Director never reads its own output — preventing the kind of runaway loops that plagued early prototypes.
Latency budget was a constant pressure. Gemini Live streaming is fast, but the total round-trip from a viewer typing a message to the host responding audibly had to stay under a few seconds to feel natural. The tick-based Director added up to 5 seconds of intentional batching delay — we had to tune this carefully to balance cost efficiency against perceived responsiveness.
Twitch OAuth2 token refresh in a long-running async process required careful lifecycle management. Tokens expire mid-stream, and the channel gateway had to handle refresh transparently without dropping IRC events or losing RTMP credentials.
Accomplishments that we're proud of
Zero-interruption mode switching — The dual-process FFmpeg architecture is genuinely novel and produces seamless transitions that are invisible to viewers. This took the most engineering effort and we're proud it works reliably.
Two-layer AI design — Separating the fast, cheap Director from the expressive, real-time Host is an elegant architectural pattern. The Director can think strategically (which product to pitch, which viewer to address) while the Host focuses entirely on sounding natural. Neither model tries to do both jobs.
Tool calling in live voice — The AI host can fetch product prices, retrieve payment links, look up viewer history, advance the product lineup, swap the on-screen visual, and post replies in chat — all mid-response, without breaking conversational flow. A viewer asks about price, the host answers aloud and posts the purchase link in chat simultaneously.
Platform-agnostic design — The Channel Gateway abstraction means the entire system above it is decoupled from Twitch specifics. Adding YouTube Live, Instagram, or TikTok is a matter of implementing one interface class, not redesigning the system.
Clean layered architecture — Every layer has a single responsibility, communicates through typed Pydantic schemas, and can be tested or replaced independently. The codebase is easy to navigate and extend.
What we learned
Broadcast-grade reliability requires thinking like a hardware engineer. Software engineers instinctively reach for "restart the process" when something goes wrong. In a live stream context, that's not acceptable. We had to adopt the mental model of a live TV director: transitions must be pre-planned, pre-warmed, and executed without gaps.
Event batching is often better than event-driven. Our first instinct was to react to every viewer event immediately. The tick-based approach felt slower conceptually, but produced far better AI responses because the Director could see the full picture of what was happening in chat before deciding what to say.
Real-time AI voice is closer to production-ready than we expected. Gemini Live's streaming audio quality and latency were genuinely good enough for a live shopping context. The main engineering work was in the surrounding infrastructure — delivery, sync, and integration — not in the AI model itself.
Async Python is excellent for this class of problem. The entire pipeline — IRC events, AI inference, audio streaming, media encoding — is I/O-bound and concurrent. FastAPI + asyncio handled it cleanly without threads or multiprocessing (except for the intentional FFmpeg subprocess design).
Pydantic at every boundary pays dividends. Enforcing typed schemas at every layer interface caught several integration bugs early and made the system far easier to reason about.
What's next for Live Shopping Agent
Multi-platform support — YouTube Live, Instagram Live, Facebook Live, and TikTok LIVE are all planned. The Channel Gateway interface is ready; it's a matter of implementing each platform's authentication and chat protocol.
AI video avatar — Replace the static product image overlay with a photorealistic AI avatar that lip-syncs to the generated audio in real time, creating a true "virtual human" shopping host.
Viewer personalization — Use purchase history and interaction patterns to tailor pitches to individual viewers. A viewer who has bought fitness products before should hear different framing than a first-time visitor.
Analytics dashboard — Real-time visibility into viewer count, engagement rate, conversion events, and which products generated the most interaction — all surfaced through the existing FastAPI/PostgreSQL backend.
Human co-host handoff — The HumanLiveContentUnit (currently planned) would allow a real presenter to seamlessly take over from the AI host when desired, with viewer context automatically handed off so the human can pick up the conversation naturally.
Campaign management UI — A web interface for brands to upload products, set pricing, configure the host persona and director strategy, and schedule campaigns — making the system accessible to non-technical users.
Self-improving Director — Log which Director instructions produced the highest engagement and use that signal to fine-tune the Director's decision-making over time.
Built With
- alembic
- async-sqlalchemy
- fastapi
- ffmpeg
- gemini-2.5-flash
- gemini-live-api
- postgresql
- postgresql-+-async-sqlalchemy-+-alembic
- python-3.11
- twitch-irc
Log in or sign up for Devpost to join the conversation.