Inspiration
Someone asked r/workingmoms, almost word for word, the question this build answers: "Is home admin (insurance renewals, boiler servicing, chasing tradespeople) actually a big pain point for people, or is it..." A few posts along, another: "Has anyone used Faye home advisor to relieve the mental load?"
They are asking because somebody is already selling this by hand. HeyFaye charges a monthly fee to do exactly this list of tasks, by people, in three Florida towns. Yohana tried it and is gone. The BLS American Time Use Survey puts household management and purchasing goods and services together at 0.86 hours a day, about 314 hours a year, and the FTC took $8.5 million from Care.com over how it sold this category.
The gap is narrower and more interesting than "household admin is hard". Every Alexa+ service provider interface Amazon publishes carries one end-user identifier, one linked account and one per-user appointment datastore. Nobody has specified the person acting on somebody else's behalf.
What it does
An adult child books the plumber into her mother's house. Her mother asks her own Echo when he is coming and who arranged it, says she needs to move it, and picks from times the plumber actually has. Both homes get a written confirmation. The mother can ask what has been changed at her house, and can pause or end the whole arrangement from her own speaker without going through anybody.
How we built it
This is a real MCP server, not a simulation of one. Streamable HTTP on @modelcontextprotocol/sdk, protocol revision 2025-11-25, 21 tools, stateless, with a fresh McpServer and transport built per request and both closed on response close, because the account a turn belongs to arrives in the Authorization header. tests/mcp-protocol.test.ts drives it with a real MCP client: handshake, tool listing, annotations, structured content, two accounts seeing two different households, and a refusal arriving as an error result rather than a protocol failure. Account linking is OAuth 2.1 with PKCE S256, the plain method refused, mismatched verifiers refused, single-use codes and rotating refresh tokens, and the web console is a real client of it rather than a bypass.
Three of Amazon's own constraints are enforced by tests rather than by intention, and that is the part of this build we would point a judge at first.
The toolkit wants a sub-500ms round trip, so tests/no-model-on-read-path.test.ts walks the import graph from the MCP server and fails if anything under src/bedrock/ or any @aws-sdk/ import ever becomes reachable from a tool handler. The latency test fails if the worst p95 exceeds a fifth of the budget. Measured worst p95, including per-request server construction, over eight tools: 6.3 milliseconds, on request_change, 1.3% of the budget.
You cannot script what Alexa says, because the model composes the sentence from the data you return. So On Behalf guarantees the fact set instead. src/domain/readback.ts enforces that every mandatory booking fact is present, that each fact is true standing alone in any order, that no field reads as an instruction to the model, and that no identifier reaches the spoken channel. Every voice result passes through it.
Voice can never be the only path, by Amazon's own accessibility rule. tests/touch-parity.test.ts walks the tool registry against the running console and fails if a tool exists that hands cannot reach, if a page carries a script tag, or if a page is reachable only by typing a URL. The console is nine routes, server-rendered, with zero JavaScript.
An add-on cannot speak first either, so nothing here announces; what_is_waiting exists precisely because nothing can come and find you. Amazon Bedrock does three jobs, all after a turn has already been answered and none reachable from the voice path: drafting an arrangement from a typed sentence, arguing a held request so the person deciding has something to read, and writing the handover note. The model never decides, never books and never cancels, and a test approves something the model recommended declining. src/domain/screening.ts refuses health content at the boundary rather than in a README, including the false positives, so "scandinavian" must not trip "scan". The console vendors two Google fonts rather than naming a system serif that would have rendered as Georgia on most judges' machines: Newsreader reserved narrowly for a person's own name and place, Source Sans 3 for everything else — labels, the ledger, the authority band. The palette is a deep pine green ground with a cold kingfisher light and one warm tone, a soft wheat, reserved for whatever is waiting on a person. 294 tests pass, 9 skipped, once with every Bedrock call forced down its fallback path and again with Bedrock live.
Challenges we ran into
The number we were about to submit was wrong, and it was our own benchmark's fault, not a rounding choice. An earlier run had reported 4.4ms worst p95 and we were ready to call the round trip "under five milliseconds" in both the README and the narration. Re-running the same harness on a fully idle machine, no emulator competing for CPU, gave 6.0 to 6.5ms consistently, with request_change the slowest tool at 6.3ms. The first number simply did not reproduce. We replaced the README's table with the real run, re-recorded the line that quoted the old figure, and left a note not to swap it back for the better-sounding one. It is still 1.3% of the 500ms budget either way, which is the number that actually matters.
The import-graph test passing was not the same thing as the model being kept off the voice path, and it took a real near miss to see the gap. A Bedrock-composed weekly brief was written to SQLite; the latest_brief tool later read that row back and handed it to the speaker's composer. tests/no-model-on-read-path.test.ts was satisfied throughout, because no tool handler ever imported src/bedrock/ or an AWS SDK module to produce that path — the model's words reached the voice channel through the database instead, one hop the import graph does not see. The health-content screen, assertNoHealthContent, was only being called on text arriving fresh from a typed request, not on text arriving from storage, so a stored brief could have carried the one category the product refuses to touch and nothing would have stopped it. The fix moved the screen into the shared screenResult wrapper every tool return passes through in registry.ts, so it runs on a fact's way out regardless of where the fact came from. The import-graph test still matters; it just is not the whole guarantee by itself.
The track mandates a protocol revision and the SDK does not tell you which one it speaks. LATEST_PROTOCOL_VERSION is a constant inside a compiled dist file, not in the README or the package metadata, so we read it out of dist/esm/types.js and then pinned a test that asserts the negotiated version on a live initialize handshake, so a dependency bump that drops the revision fails the suite rather than failing certification.
The SDK's own type declarations point you at the deprecated API: six tool(...) overloads appear first, each marked "use registerTool instead", with registerTool below all of them. And nothing in the Streamable HTTP documentation describes either mode in terms of where per-request credentials go, which is the first question anyone building against account linking has, so the working pattern had to be assembled from type declarations.
Node's own TypeScript runtime and tsc disagree about what TypeScript is. Constructor parameter properties are not erasable syntax, so six error classes type-checked clean and threw at module load. erasableSyntaxOnly is off by default while running a .ts file directly is on by default, which puts the failure at runtime in the one place a type checker was meant to cover.
Accomplishments that we're proud of
npm run demo runs the whole story end to end over the real protocol, two MCP clients with two different bearer tokens, nothing reached into the database to make anything true. It was verified twice, once with Bedrock live and once with the model switched off entirely.
What we learned
Three of the constraints that read as limitations are really a specification. Once "cannot speak first", "cannot script the sentence" and "voice cannot be the only path" were written as tests rather than as notes, most of the design questions answered themselves.
What's next
A real mail and SMS provider behind the written channel, adjudication written in the background so the delegate never sees "not yet reviewed", and the same delegation model offered to the service provider interfaces that currently assume one person per account.
Built With
- amazon-bedrock
- aws-sdk
- model-context-protocol
- node.js
- typescript
- zod

Log in or sign up for Devpost to join the conversation.