Inspiration
Alexa+ is stateless by default. You ask a question, get an answer, and five minutes later that context is gone. If you ask a follow-up, you have to repeat yourself. This is a real friction point when your hands and eyes are busy — cooking, working, driving.
I wanted to see if I could give Alexa+ something it does not naturally have: persistent context across separate interactions. Not a chatbot with memory bolted on, but a proper MCP server that exposes context tools the agent can call when it needs them. When Amazon opened the Alexa+ track for this hackathon, it was the perfect opportunity to build it against the real protocol.
What it does
Context Keeper is a self-hosted MCP server (spec 2025-11-25+, Streamable HTTP) that maintains a lightweight, per-user context thread across Alexa+ interactions.
It exposes three tools:
start_context(user_id, topic)— begins a new context threadquery_context(user_id)— retrieves the current context threadlist_contexts()— lists all active threads
When you ask a question, the agent calls query_context first. It gets back what you were just doing, then answers with that context folded in. Ask "what was that temperature again?" five minutes after asking about a roast chicken, and it knows what you mean.
How I built it
Stack: Python, FastMCP, Streamable HTTP transport, in-memory store.
The core is the MCP server. It implements the Streamable HTTP transport, registers the three tools above via the @mcp.tool() decorator, and handles per-user session state. A second Python script acts as the MCP client — it speaks the same protocol Alexa+ would use, so the demo is a real end-to-end MCP exchange, not a mock.
I chose Python because the FastMCP SDK makes tool registration declarative. Registering a function with @mcp.tool() exposes it over the protocol without writing any custom JSON-RPC handling. The server starts in seconds with uv run server.py and listens on http://127.0.0.1:8000/mcp.
Challenges I faced
- Spec versioning: The hackathon requires spec 2025-11-25 or later with Streamable HTTP. I confirmed FastMCP 3.4.7 uses the current spec and exposes the right transport.
- Statelessness vs. context: MCP servers are designed to be stateless by default. Maintaining per-user context meant designing a session model that keeps the protocol contract intact.
- Demo without hardware: Since I do not have a physical Alexa+ device, I built a client that exercises the same code path the real agent would. This made the demo clearer, because the judge can see the MCP traffic directly in the server logs.
What I learned
MCP is a clean, minimal protocol. The tool-registration model — decorate a function, get a networked tool — is a strong pattern for exposing capabilities to LLM agents. The Streamable HTTP transport is straightforward once the SDK handles it. I also learned how important it is to demonstrate real protocol traffic, not just describe it, when the judges cannot see the code run.
What's next
- Persistent storage so context survives server restarts
- Context expiration and privacy controls
- Multi-user support with isolated threads
Log in or sign up for Devpost to join the conversation.