Inspiration

Online shopping assumes you can see. Not just the checkout, the whole thing: you skim twenty results in two seconds, glance at a cart badge, scan a review page before paying. Take sight away and every one of those becomes a sentence somebody has to read to you, in order, with no way to skip ahead.

Screen readers make a shop operable. They do not make it quick.

That is the whole premise, so I stopped asserting it and measured it. I opened seven live WooCommerce shops in Chrome and walked Chrome's own accessibility tree over CDP, the same tree NVDA, JAWS and VoiceOver are handed, counting every announcement until the first product is named.

The median category page spends 98 words before it names a single product. One spends 270. Jumping straight to the main landmark, which is what an experienced screen reader user actually does, still costs a median of 55. Comparing three products, the thing a sighted shopper does in one glance, costs 93 words.

VoiceCart names the first product in word one, compares three in 40 words, and finishes an entire order in 119. That last number is the one I would put in front of a judge: a whole order here costs less than half of what one of those shops spends getting to its first product.

Both scripts are in the repo and run against any shop on the web in under a minute, because a number you cannot check is just a louder assertion.

What it does

VoiceCart is an MCP server a voice assistant can shop. You browse departments, hear products described, fill a basket, order, and check where the order has reached, without a screen ever being involved.

Everything about it is shaped by how hearing works rather than how a page looks.

Three at a time, never twenty. A listener cannot skim, so results arrive three at a time with a count of what is left, and the assistant offers rather than continues.

The disqualifying fact first. "Out of stock" is said before the price, not after it, so nobody spends a sentence deciding to buy something they cannot have.

"The second one", because nobody says a SKU. On a screen you point; in a conversation you refer back. The server remembers the last list it read you and resolves against it, so positions, names and a bare "that one" all work. The recent list beats the rest of the shop: two products have "cotton" in the name, but right after hearing one of them, "add some cotton" is not ambiguous. When it genuinely is, nothing is added and the reply asks which.

Speech corrects itself. Nobody shops out loud in clean commands. They say "no, not that one", "take the last one back", "start again", halfway through the sentence they are already in. So a correction is its own intent. Pass the shopper's words to repair and the shop walks backwards: it restores the basket, and "not that one" additionally stops that product being offered for those same words again, because that is what they were telling you.

The undo restores a snapshot of the whole basket rather than reversing an operation, because an add is not reversible on its own; it may have been clamped to what was in stock, or merged into a line already there. Only a placed order cannot be undone. Handing somebody back a basket they have already bought would be a lie about what they owe.

A basket that outlives the conversation. A sighted shopper leaves a browser tab open; a voice shopper has no tab. Baskets are keyed to the shopper and persisted, which is also what makes "order the usual again" mean something.

One sentence is the review page. There is nothing to glance at before paying, so the confirmation carries the two facts that cost money to get wrong, the amount and the address, and nothing else.

Things you would have seen, said instead. Allergens on food, every time. A warning when the parcel is heavy enough that somebody should be at the door. Low stock when there are three left.

The shop is cash on delivery, which is how most of South Asia buys online. It also removes the worst moment in voice commerce: nobody is ever asked to say a card number out loud, in a room, to a device.

How I built it

Python and the MCP Python SDK, served over Streamable HTTP. The server negotiates protocol 2025-11-25.

Twelve tools, each returning the same declared shape: a speech string to read aloud exactly as written, and a cards list to render only if there is a screen. Declaring that in the output schema rather than describing it in prose is what lets a client show a product carousel instead of reading JSON.

Beyond tools, the rest of the protocol is used where the rest of the protocol is the right answer:

Resources, because reading is not an action. shop://catalogue, shop://category/{name}, shop://cart/{shopper} and shop://order/{id} are side-effect free and addressable.

Subscriptions, because a resource that has to be re-read to be trusted is only half a resource. A client watching a basket is told the moment it changes, rather than polling to find out.

Completion, so the assistant offers a department that exists rather than guessing at one.

A prompt, shop_by_voice, telling a client how to run the conversation.

Elicitation, the interesting one. place_order does not trust the assistant's judgement. If the client can ask, the shop puts the order to the shopper through the protocol and waits for the answer. Elicitation is optional in MCP, so it degrades to handing the question back rather than failing. Either way nothing is ordered without a yes.

It runs against live WooCommerce with four settings and no code changes. Products are read from the shop, orders are written back as real cash-on-delivery orders. The work there is the mapping, not the HTTP: a WooCommerce payload is written for a page, so it arrives with HTML in the description, a price of "" on variable parents, stock_quantity: null where stock management is off, and no allergen field at all. A page can afford that vagueness; a listener cannot skim past it.

Seventy five tests. Four of them start the server as its own process and drive it through a real MCP client over Streamable HTTP, including an order confirmed by elicitation, one declined, and a client subscribing to a basket and being told when it changes. Sixteen cover the WooCommerce mapping against a stdlib fake shop, so they run offline.

The demo video is rendered from a live run of the server rather than recorded, so it cannot drift from what the code actually does.

Challenges I ran into

The pricing of ambiguity. My first resolver matched substrings, so "hon" found honey, which is convenient right up to the moment it hands somebody the wrong jar. Matching whole spoken words made "cotton" ambiguous across two products, and the honest fix was not a better guess but refusing to guess: add nothing, ask which. For somebody who cannot see the basket, a wrong item is worse than a second question.

Elicitation took a rethink too. The obvious use is asking for the address. The better one turned out to be the confirmation, because that is the moment where an assistant's own judgement should not be enough.

Subscriptions cost me the most, and both problems failed silently. The SDK's high-level server never registers a resources/subscribe handler, and at 2025-11-25, the version this negotiates, the advertised capability is derived from whether that handler exists. So the server was honestly telling clients it could not be subscribed to, while the only notify API it offered published to a stream from a later protocol version that nobody was listening on. Then, having registered the handler, every notification still vanished: over Streamable HTTP a fresh ServerSession is built per request, so the session object that subscribes is not the one that later changes the basket. Keying subscribers on it looks exactly like working code until you check whether anything arrived. The transport's own Mcp-Session-Id is the thing that survives. Both are written up in the friction log with reproductions.

Accomplishments that I am proud of

That the safety is in the shape rather than in warnings. place_order cannot place an order without a yes, whichever route the question travels. add_to_cart cannot oversell, because quantity is clamped to real stock. An ambiguous phrase cannot put anything in the basket. A placed order cannot be undone. None of that depends on the assistant behaving well.

And that the premise is measured. It would have been easy to keep asserting that screen readers are slow. Counting it turned the pitch into two numbers anybody can reproduce.

What I learned

That "accessible" and "voice-enabled" are not the same claim. Reading a screen-shaped API aloud is voice-enabled and still exhausting. The work was in the decisions a page makes silently: how much to say at once, what order to say it in, and what a listener is owed that a reader can simply look at.

And that a conversation is not a form. A form only moves forward. Real speech spends a good part of its time taking things back, and until I built repair, every correction was work the shopper had to do on my behalf.

What's next for VoiceCart

Learning a shopper's usual basket rather than only their last order, and a second transport so the same shop works from a phone as well as a speaker.

Built With

  • alexa
  • mcp
  • model-context-protocol
  • pydantic
  • pytest
  • python
  • rest-api
  • streamable-http
  • uvicorn
  • woocommerce
Share this project:

Updates

Submission history