Inspiration
Every evening my family asks the same question: did anyone come to the door today? Finding out means opening a vendor NVR app and scrubbing through hours of footage. Most homes and small shops in India run a mix of Hikvision, Dahua, Frigate, Synology or ONVIF cameras, and none of them can simply be asked. I wanted Alexa+ to answer that question without the video ever leaving the house.
What it does
cctvQL for Alexa+ is a self-hosted MCP server (spec 2025-11-25, Streamable HTTP) that lets Alexa+, or any MCP client, answer plain questions about the cameras people already own:
- "Was anyone at the front door today?"
- "Did a package arrive?"
- "What happened in the last hour?"
- "Are all my cameras online?"
- On an Echo Show: "Show me the driveway."
Every answer is written to be spoken aloud, for example "Yes. 2 people detected at the Front Door in the last day. The latest was 50 minutes ago.", and it comes with structured data for screens.
How we built it
cctvQL is my existing open source (MIT) project: a conversational layer that normalises events from Frigate, Hikvision, Dahua, Synology Surveillance Station, Milestone XProtect, Scrypted and ONVIF behind one adapter interface. For this hackathon I built the Alexa+ side with the official MCP Python SDK (FastMCP, stateless Streamable HTTP):
- Seven voice-first tools:
list_cameras,recent_activity,who_was_at,latest_snapshot,camera_health,ask_cameras(free-form, through cctvQL's own query engine), and opt-inpoint_camerafor PTZ presets - Speech-friendly answers: relative times ("50 minutes ago"), plurals, aliases (someone to person, parcel to package), and loose spoken-name matching ("front door" finds
Front_Door) - Security: bearer-token auth (
CCTVQL_MCP_TOKEN), with a warning if the server is bound beyond localhost without a token cctvql mcpone-command launcher, plus a live demo adapter so anyone can try it without cameras- End-to-end tests that drive the server with the official MCP client over real Streamable HTTP; the full suite (428 tests) passes in CI on Python 3.10 to 3.13
Privacy by design
The server runs on the home network next to the recorder. Only short text answers and snapshot links leave the house, never video frames, and without the token the server returns HTTP 401.
Challenges we ran into
- Writing answers that sound natural when read aloud, rather than JSON dumps
- Matching how people say camera names to how installers name them
- MCP Python SDK 2.x arrived with breaking client changes mid-build; cctvQL pins the tested 1.x line for now
What's next
- Test inside Alexa+ Preview as access allows, and tune answer length for voice
- Proactive announcements ("a package just arrived") via cctvQL's existing alert engine
- Caretaking mode: a daily "all quiet" summary for family members of elderly people living alone
Built during the hackathon
cctvQL existed before (v1.0.4 on PyPI). Everything for Alexa+ is new in this window: the MCP server and tools, the cctvql mcp command, token auth, the live demo adapter, docs and the end-to-end tests. See pull request #8.
The demo uses cctvQL's built-in demo cameras; every answer shown is a real response from the running server. The voiceover is synthetic.
Log in or sign up for Devpost to join the conversation.