Cue AI

Inspiration

Millions of people can see a screen and speak clearly but cannot reliably operate a mouse, a trackpad, or a multi-step checkout form. This can be due to tremors, limited hand mobility, fatigue, or just not having a free hand. Online shopping assumes you have all three. We wanted to give back the part of shopping that never needed hands in the first place, by telling someone what you want and having them go find it, compare it against the other one, and hand it to you, while never once reaching into your wallet without asking first.

What it does

Cue shops a real site like Amazon and H&M by voice and gaze. You say or look at what you want, and Cue reads the actual page. You can ask it what's the highest rated one under a hundred dollars, and it answers from the page in front of it.

Ask it to compare two products and a panel opens over the page: a plain-English verdict first, then the two products side by side with their real prices, then what buyers actually said about each, a direct quote if the page genuinely has one, and if it doesn't (which, we learned, is now the common case on Amazon, more on that below), an honest reading of the star breakdown instead: "most people love it, but one in six gave it one star." If you've bought something recently, Cue will connect the two out loud: "you picked up an iPhone last week, these pair the moment you open the case", but only when that's actually true.

How we built it

The client is a Manifest V3 Chrome extension with no framework, a single event bus that gaze, voice, the overlay, and the avatar all publish to and subscribe from. It injects into live third-party pages, so Cue works on a website's actual DOM rather than a page we control. Product and page-control discovery walks the accessibility tree (role, aria-label, tag semantics) — the same information every real site already exposes for screen readers — rather than depending on site-specific selectors that would break the moment a layout changed.

Speech runs through an AudioWorklet capturing 16kHz PCM16, streamed over our own WebSocket to Grok's speech-to-text, and the browser's built-in SpeechRecognition needs near-silence and drops the connection every ~60 seconds. Text-to-speech falls through Grok → ElevenLabs → the browser's own voice, with every line disk-cached so a rehearsed demo costs nothing to repeat.

The reasoning layer is Grok API, called once per turn with a narrow, explicit action vocabulary and a strict JSON contract, like search, compare, open a link, scroll, add to cart, and so on. A hand-written regex router in front of it (server/router.py) catches the exact phrases that must never wait on a model round trip and must never be something a model could get creative with: "yes," "stop," "no thanks" to a popup. That router is also the only place the words that actually commit money approve_checkout are ever allowed to originate.

Underneath all of it is a full gaze-tracking pipeline WebGazer.js with a TensorFlow.js face mesh, wrapped in our own outlier gate, One Euro filter, and a from-scratch 13-point calibration flow.

Challenges we ran into

Webcam gaze is worse than advertised, and that reshaped the whole product. We measured 220–350px of real error against WebGazer's own published figures, which is wider than an Amazon product tile. A pointer that's off by more than the thing it's pointing at isn't a pointer, it's a hint. WebGazer also silently retrains itself on your mouse clicks by default, which was quietly poisoning our accuracy numbers until we found and removed those listeners.

Amazon doesn't ship review text to a signed-out browser anymore. We went looking for a stale CSS selector and found something more interesting: fetched a live listing's full HTML and DOM, searched for every review-text hook we could find, and got zero matches on all of them. The reviews section renders only a star-percentage histogram now. Rather than fail quietly, Cue reads that histogram and says the actual split out loud, and because we verify every review "quote" the model produces against the page's own text before it's ever shown, a fabricated quote is structurally impossible, not just discouraged by a prompt.

The extraction layer was inventing products out of page furniture. Reading the whole page instead of just the viewport surfaced a bug that had been mostly hidden before: a generic "big linked box" heuristic was tagging things like "+1 other color/pattern" and "29,222 ratings" as real products, and the agent would occasionally discuss them as if they were items on the page. This was fixed by requiring actual evidence, like a real product-shaped URL or a real price, before anything counts as a product at all.

Accomplishments that we're proud of

Cue placed a real order on real Amazon by voice, start to finish with no code path that could have skipped the spoken "yes."

The trust boundary holds at two independent layers, not one. The model is physically incapable of producing confirm, approve_checkout, or setup_passkey, those verbs are stripped from its output before the client ever sees them. And the client's own gate is keyed on whether the instruction came from the deterministic router, not on any model's name or wording, so replacing our reasoning model mid-project (we started on Claude, switched to Grok) couldn't have silently reopened anything, because the gate was never about the model in the first place.

Cue is honest about what it doesn't know. When it can't place which item you mean, it says what it can actually see instead of guessing. When a comparison has no real review text to draw from, it says so and reasons from what it does have rather than inventing a quote.

What we learned

Measure your own numbers, always, on the real thing. One afternoon of measuring gaze accuracy against a live face changed every subsequent design decision, because it turned gaze from "the pointer" into "a rough region" and made voice the thing that actually commits. The same discipline caught the review-hook problem, the fake-product-extraction bug, and a caching cap that was silently discarding the last third of every fetched page, which were all found by fetching the real page and counting bytes.

We also learned how much of accessibility is tone, not just capability. Reading an item's name, price, and what people actually said about it back out loud, in the words a person would use, not a spec sheet, is sometimes the only description a shopper gets before agreeing to buy something. Getting that voice right mattered as much as getting the checkout logic right.

What's next for Cue

Delivering the signed approval to a second device. The signing, the spending caps, and the exact-words record are already built and live; today, approval is a spoken "yes" plus a passkey on the same device the shopper is using. The next step is routing that same signed request to a parent, a carer, or a family member's own phone, so the person approving a purchase doesn't have to be the person who can't reach the checkout button.

Built With

Share this project:

Updates

Submission history