Inspiration

Around forty million people worldwide live with aphasia, most after a stroke. The person is entirely there — comprehension, intelligence, opinions — but the sentence gets stuck on the way out. A phone call is the worst case: no gestures, no face, a receptionist who will hang up. Many simply stop making calls and hand their independence to whoever will make them instead.

Existing tools either need weeks of training recordings before they work, or hand the user a keyboard. We wanted something that works from the first syllable — and that never puts words in someone's mouth without their say-so.

What it does

Relay listens to whatever does come out — a syllable, a broken word, a tapped concept — and shows three or four complete things the person might be trying to say. They tap one. It is spoken aloud in a natural voice. If the sentence asked for something to happen, a Strands agent does it.

  • Live call view. A simulated receptionist, pharmacy, or family member is on the line. When they ask a question, replies appear before the user has said anything, then sharpen as fragments come in. Tap, and it's spoken.
  • Telegram. Family message or voice-note the bot; the user gets reply candidates and taps one; it's delivered. /say turns the phone into her voice in the room: speak, tap, the phone says it. In any chat, @relay1517_bot om tues pops up the candidates and sends the chosen one as her own message.
  • Actions that really happen. "Tell Sam I'll be late" reaches Sam's phone. "Send my scan to Dr. Chen" lands a PDF in his chat. "Call Dr. Chen for an appointment" asks him to ring her. "Remind me in two minutes" buzzes her phone two minutes later — the agent computes the time, the server fires it.
  • Emergencies. "Help… fire…" is detected in code, not left to the model: every adult contact gets her exact words and a map pin (shared with her consent through Telegram's own location prompt), and the fire brigade or ambulance is called. It works even if the AI provider is down.
  • It learns. Every chosen sentence is remembered and recalled by sound — a mumbled "om" brings back "I want to go home". Share a contact card from her phone book and Relay knows that person; /note my neighbour Ruth checks on me on Fridays and it knows that too. Only what she chooses to give it — it never reads other chats, call logs, or notes.
  • It never goes silent. If the model is slow or throttled, an offline engine answers inside the deadline, with her own words offered verbatim if nothing else fits.

How we built it

Strands Agents, in two deliberately different shapes.

The prediction loop is a Strands Agent with no tools and a Pydantic structured_output_model — one constrained call, no agent loop. Someone is waiting on the line; a correct suggestion that lands two seconds late is a failed suggestion. We force the structured-output tool from the first request and hold the call to an eight-second deadline before falling back offline.

The action executor is the opposite: a real Strands agent with @tool functions — send_message, send_document, set_reminder, place_call, share_location, alert_emergency, order_item — and a genuine tool loop. It runs once, after the user commits. We watched it chain send_document and set_reminder from a single sentence without being told to. Tools read the session from Strands' ToolContext.

A third small Strands agent plays the person on the other end of the call.

Around them: Python/FastAPI; React + Vite PWA; Groq Whisper for speech in, with a vocabulary hint from her profile and a confidence filter for hallucinations; neural voices for speech out; the browser captures raw audio with a 350 ms pre-roll so the soft "h" in "home" is never clipped. Model providers are a config choice — Bedrock, Anthropic, OpenAI, Groq — because Strands normalises them behind one Agent. Telegram via long-polling, no public URL. Personalisation via fuzzy sound-matching over every past choice, with an Amazon Bedrock AgentCore Memory backend behind the same interface.

Challenges we ran into

  • The model fought the user. Told the fragments were "toothpaste, toothpaste", it kept suggesting clinic appointments — because that's what callers to a clinic usually want. Fixing that took a prompt rule ("her words win"), a code guard that keeps her clearly-spoken words on screen, and stripping unrelated past choices out of the prompt: a sentence she chose yesterday was leaking into today's suggestions as a fact.
  • Never invent. The first live model turned "tues" into "Tuesday at 3pm". Our top rule is never invent facts, backed by a guard that offers her words verbatim if the model drops them.
  • Speech recognition autocorrects the wrong way. The browser's recogniser turned broken speech into fluent, wrong words. Whisper with a glossary keeps fragments as fragments — but hallucinates on silence, so we filter on its own confidence scores.
  • Free-tier limits are demo killers. A throttled provider was being retried six times with backoff — 130 seconds with someone on the line. Now: no retries on the hot path, a hard deadline, a circuit breaker to a second model, and an offline engine underneath. An emergency alert runs the tools directly if the model is unavailable.

Accomplishments that we're proud of

  • Replies appear before the user speaks, from the other side's question alone.
  • The two-shape Strands design: a single constrained call where latency is the feature, a real tool-using agent where deliberation is.
  • A sentence chosen by one tap arrives on a family member's phone; three broken words send a document to the doctor; "fire… help" alerts everyone who can act, with a map pin, in one tap.
  • Zero enrolment. It works on the first sentence.
  • It degrades gracefully instead of going blank — and says so on screen.

What we learned

The hard part of an assistive agent is not what it can do but what it must refuse to do: invent a time, swap a topic, speak without a tap. Most of our engineering went into guards around the model, not the model. And Strands lets you choose the shape of each call — a structured single shot or a full agent loop — which is exactly the choice a latency-critical product needs.

What's next for Relay

  • Real telephony (Twilio media streams) — the reply engine is already built for it.
  • Claude Haiku on Bedrock for the prediction tier, and AgentCore Memory in production so the phrasebook follows the user across devices.
  • Onboarding for families and speech therapists to build the profile that makes "chuh" resolve to "Dr. Chen".
  • A native app, so — with permission — contacts and call history can inform suggestions too.

Built With

  • amazon-bedrock-agentcore
  • edge-tts
  • fastapi
  • groq
  • pydantic
  • python
  • react
  • strands-agents
  • telegram-bot-api
  • typescript
  • uvicorn
  • vite
  • web-speech-api
  • whisper
Share this project:

Updates

Submission history