Inspiration

Every morning starts the same way: I wake up and the first thing I do is grab my phone. One notification becomes ten, and before I have brushed my teeth I am already deep in apps, messages and tasks. I wanted my mornings back. What if I could get things done while I am cooking or brushing my teeth, just by talking, and get updates and details without ever opening my phone? Even for basic stuff, I would rather just ask Alexa than open yet another app.

I already run a personal AI agent called Hermes. It has skills, memory and a Discord gateway. The missing piece was a voice front door. Alexa+ talks to remote MCP servers, but Hermes only ships a local stdio MCP server, so I built the bridge.

What it does

Hermes Home is a self-hosted MCP server (Streamable HTTP, spec 2025-11-25) that lets Alexa+ drive my Hermes agent by voice:

  • Instant replies, work in the background. Alexa answers right away ("Let me ask Hermes", "On it"). Slow work runs asynchronously, so a voice assistant never waits on an agent.
  • Hands-free actions. "Send me a message on Discord with a question to think about" and Hermes writes it and delivers it through its own gateway. "Make a Homeless Entrepreneur carousel and send it to my Discord" and Hermes runs its carousel skill, renders eight slides and delivers them.
  • Alexa handles everyday chat herself; Hermes is only used when I say "Hermes" or ask for an action.
  • 14 hand-picked tools (never the whole agent catalogue), plus optional calendar, Home Assistant and GitHub tools.

How I built it

  • TypeScript, Fastify and the official MCP SDK for the Streamable HTTP server (one McpServer per session).
  • Hermes Agent API (/v1/chat/completions and /v1/runs) on a cloud VM, reached over an SSH tunnel, running Z.ai GLM-5.3-Flash.
  • Alexa+ simulator. Alexa+ is in Preview and I do not have access, so I built a simulated Alexa+ web app (voice and chat) that acts as the MCP client using the official SDK. It discovers tools, calls them over Streamable HTTP and shows every tools/call in a live trace panel.
  • Safety by design: fail-closed Hermes gate, per-user isolation of runs and memory, DNS-rebind guard, two-step confirmation for home control, RS256/JWKS OAuth2 mode for internet exposure, and redaction of private identifiers in the simulator trace.

Challenges

  • Voice needs sub-second responses, while a real agent takes 15 to 90 seconds. The async-by-default design and instant acknowledgements solve that.
  • Model providers cost the most time: free tiers are too small for agent contexts, and a Z.ai Coding Plan key only works on its own endpoint.
  • The MCP SDK allows one transport per server instance, so every session gets its own server.
  • Hermes skills with owner allowlists correctly refuse API callers. I did not bypass that. (Full details are in the friction log.)

What I learned

Treating the agent as the brain and the voice assistant as a thin front door works well. MCP makes that clean, as long as you design for latency and keep the tool surface small.

What is next

Onboard to real Alexa+ once Preview access opens, add calendar and Home Assistant for morning briefings, and let Hermes proactively read me my day while I make breakfast.

Built With

Share this project:

Updates

Submission history