Inspiration

Alexa+ is stateless by default. You ask a question, get an answer, and five minutes later that context is gone. If you ask a follow-up, you have to repeat yourself. This is a real friction point when your hands and eyes are busy — cooking, working, driving.

I wanted to see if I could give Alexa+ something it does not naturally have: persistent context across separate interactions. Not a chatbot with memory bolted on, but a proper MCP server that exposes context tools the agent can call when it needs them. When Amazon opened the Alexa+ track for this hackathon, it was the perfect opportunity to build it against the real protocol.

What it does

Context Keeper is a self-hosted MCP server (spec 2025-11-25+, Streamable HTTP) that maintains a lightweight, per-user context thread across Alexa+ interactions.

It exposes three tools:

  • start_context(user_id, topic) — begins a new context thread
  • query_context(user_id) — retrieves the current context thread
  • list_contexts() — lists all active threads

When you ask a question, the agent calls query_context first. It gets back what you were just doing, then answers with that context folded in. Ask "what was that temperature again?" five minutes after asking about a roast chicken, and it knows what you mean.

How I built it

Stack: Python, FastMCP, Streamable HTTP transport, in-memory store.

The core is the MCP server. It implements the Streamable HTTP transport, registers the three tools above via the @mcp.tool() decorator, and handles per-user session state. A second Python script acts as the MCP client — it speaks the same protocol Alexa+ would use, so the demo is a real end-to-end MCP exchange, not a mock.

I chose Python because the FastMCP SDK makes tool registration declarative. Registering a function with @mcp.tool() exposes it over the protocol without writing any custom JSON-RPC handling. The server starts in seconds with uv run server.py and listens on http://127.0.0.1:8000/mcp.

Challenges I faced

  • Spec versioning: The hackathon requires spec 2025-11-25 or later with Streamable HTTP. I confirmed FastMCP 3.4.7 uses the current spec and exposes the right transport.
  • Statelessness vs. context: MCP servers are designed to be stateless by default. Maintaining per-user context meant designing a session model that keeps the protocol contract intact.
  • Demo without hardware: Since I do not have a physical Alexa+ device, I built a client that exercises the same code path the real agent would. This made the demo clearer, because the judge can see the MCP traffic directly in the server logs.

What I learned

MCP is a clean, minimal protocol. The tool-registration model — decorate a function, get a networked tool — is a strong pattern for exposing capabilities to LLM agents. The Streamable HTTP transport is straightforward once the SDK handles it. I also learned how important it is to demonstrate real protocol traffic, not just describe it, when the judges cannot see the code run.

What's next

  • Persistent storage so context survives server restarts
  • Context expiration and privacy controls
  • Multi-user support with isolated threads

Built With

Share this project:

Updates

Submission history