Inspiration
Every year, tens of thousands of police pursuits unfold across the United States. Analyzing these events after the fact requires hours of manual video review — scrubbing through helicopter footage, estimating speeds, tracking distances, identifying vehicles. We built Pursuit Analyst to compress that entire process into 40 seconds using the full breadth of Google's Gemini ecosystem.
What it does
Give it a YouTube URL of a police chase. Click one button. In under a minute, you get:
- Semantic tactical phases — every distinct moment identified, timestamped, and scored by danger level (1–5)
- Suspect vehicle identification — make, model, year, color, damage assessment, plus an AI-generated reference image via Imagen 4
- High-frequency tracking — 50 data points mapping suspect and police speeds, distances, and environment changes
- Dual vector embeddings — every phase indexed as both text and multimodal video embeddings, enabling semantic search across what was described AND what was visually seen
- Pre-cached audio commentary — AI-generated voice analysis for every phase, played in sync with the video
- AI-generated summary video — an 8-second cinematic reconstruction via Veo 3.1
And then you can talk to it. Out loud. In real time. Ask what happened at minute 3. Ask about the suspect. Ask about the most dangerous moment. It answers instantly with full context.
How we built it
The architecture uses 7 Google AI APIs working together:
- Gemini 3 Flash — Three parallel analysis tasks (semantic chunking, suspect tracking, vehicle ID) complete in ~23 seconds
- Gemini Embedding — Each phase is embedded twice: as text (768-dim) and as multimodal video, into two separate Vertex AI Vector Search indexes
- Vertex AI Vector Search — Dual indexes with STREAM_UPDATE, cosine distance, and token restricts for filtering
- Imagen 4 — Generates a clean reference image of the suspect vehicle from the description
- Gemini Native Audio — Pre-caches all TTS commentaries AND powers the live voice agent via ADK bidi-streaming
- Google ADK — LiveRequestQueue + run_live() for real-time bidirectional voice conversation with function tools
- Veo 3.1 — Generates an 8-second cinematic pursuit reconstruction from the top danger-scored phases
The frontend is a single HTML file — no framework, no build tools. Two WebSocket endpoints handle everything: /ws/chat for analysis and /ws/voice for live audio streaming.
Deployed on Cloud Run with --no-cpu-throttling and --session-affinity for WebSocket persistence.
Challenges we ran into
- Vertex AI Vector Search doesn't store rich metadata — we built a GCS-based metadata layer mapping vector IDs back to full phase data
- Veo requires explicit file download — the
generate_videosresponse contains a reference, not bytes; must callclient.files.download()separately - Voice agent latency — solved with bounded async queues (maxsize=200), reduced context injection (top 5 phases only), and minimal tool set (2 tools instead of 4)
- Native audio context fills fast — ~25 tokens/second of audio means context management is critical for longer conversations
Accomplishments that we're proud of
- 40-second full analysis of a 10-minute pursuit video
- Dual embedding search — text and video indexes return different top results ~40% of the time
- Zero-latency audio commentary — pre-cached TTS plays perfectly synced with video
- 7 Google APIs orchestrated in a single coherent application
- Single HTML file frontend with no build tooling
What we learned
- Dual vector indexes (text + video) capture fundamentally different information and are worth the complexity
- Keep voice agent instructions concise — every token adds latency on every turn
- Cloud Run WebSockets need
--no-cpu-throttlingand--session-affinityor they break randomly - Bounded async queues with drop-oldest policy prevent audio buffering issues in real-time streaming
What's next for Pursuit Analyst
- Multi-video comparison across incidents
- Real-time live stream analysis
- Integration with law enforcement databases for automatic BOLO generation
- Multi-language support for international deployments
Built With
- fastapi
- ffmpeg
- gemini-3-flash
- gemini-embedding
- gemini-native-audio
- google-adk
- google-cloud
- google-cloud-run
- google-secret-manager
- imagen-4
- javascript
- python
- veo-3.1
- vertex-ai-vector-search
- web-audio-api
Log in or sign up for Devpost to join the conversation.