Inspiration

Every evening my family asks the same question: did anyone come to the door today? Finding out means opening a vendor NVR app and scrubbing through hours of footage. Most homes and small shops in India run a mix of Hikvision, Dahua, Frigate, Synology or ONVIF cameras, and none of them can simply be asked. I wanted Alexa+ to answer that question without the video ever leaving the house.

What it does

cctvQL for Alexa+ is a self-hosted MCP server (spec 2025-11-25, Streamable HTTP) that lets Alexa+, or any MCP client, answer plain questions about the cameras people already own:

  • "Was anyone at the front door today?"
  • "Did a package arrive?"
  • "What happened in the last hour?"
  • "Are all my cameras online?"
  • On an Echo Show: "Show me the driveway."

Every answer is written to be spoken aloud, for example "Yes. 2 people detected at the Front Door in the last day. The latest was 50 minutes ago.", and it comes with structured data for screens.

How we built it

cctvQL is my existing open source (MIT) project: a conversational layer that normalises events from Frigate, Hikvision, Dahua, Synology Surveillance Station, Milestone XProtect, Scrypted and ONVIF behind one adapter interface. For this hackathon I built the Alexa+ side with the official MCP Python SDK (FastMCP, stateless Streamable HTTP):

  • Seven voice-first tools: list_cameras, recent_activity, who_was_at, latest_snapshot, camera_health, ask_cameras (free-form, through cctvQL's own query engine), and opt-in point_camera for PTZ presets
  • Speech-friendly answers: relative times ("50 minutes ago"), plurals, aliases (someone to person, parcel to package), and loose spoken-name matching ("front door" finds Front_Door)
  • Security: bearer-token auth (CCTVQL_MCP_TOKEN), with a warning if the server is bound beyond localhost without a token
  • cctvql mcp one-command launcher, plus a live demo adapter so anyone can try it without cameras
  • End-to-end tests that drive the server with the official MCP client over real Streamable HTTP; the full suite (428 tests) passes in CI on Python 3.10 to 3.13

Privacy by design

The server runs on the home network next to the recorder. Only short text answers and snapshot links leave the house, never video frames, and without the token the server returns HTTP 401.

Challenges we ran into

  • Writing answers that sound natural when read aloud, rather than JSON dumps
  • Matching how people say camera names to how installers name them
  • MCP Python SDK 2.x arrived with breaking client changes mid-build; cctvQL pins the tested 1.x line for now

What's next

  • Test inside Alexa+ Preview as access allows, and tune answer length for voice
  • Proactive announcements ("a package just arrived") via cctvQL's existing alert engine
  • Caretaking mode: a daily "all quiet" summary for family members of elderly people living alone

Built during the hackathon

cctvQL existed before (v1.0.4 on PyPI). Everything for Alexa+ is new in this window: the MCP server and tools, the cctvql mcp command, token auth, the live demo adapter, docs and the end-to-end tests. See pull request #8.

The demo uses cctvQL's built-in demo cameras; every answer shown is a real response from the running server. The voiceover is synthetic.

Built With

Share this project:

Updates

Submission history