Inspiration

Every first meeting between two people starts the same broken way: one side already has all the context, and the other has to rebuild it from zero. A patient tells the receptionist their symptoms, then the nurse, then the doctor — third time in ten minutes, with less patience each time. A sales rep walks into call nine of the day cold and has to ask "remind me what your team needed again?" A financial advisor spends the first ten minutes of every meeting reconstructing a conversation from six months ago instead of actually advising.

Different rooms, same tax: context that already existed, thrown away and rebuilt every single time. We wanted to build something that closes that gap before the meeting even starts — not a chatbot bolted onto a calendar, but a pipeline that gathers context ahead of time and hands it straight to the people who need it.

We call the idea Context as a Service (CaaS). Healthcare is where we built and proved it first.

What it does

CaaS is an agentic pipeline that runs an entire pre-visit conversation for a patient, before they ever reach the doctor — and hands the doctor a structured report before the patient walks into the room.

  • A patient books an appointment, gets an SMS link, and taps it to attach to their session.
  • The moment they do, a fresh Strands BidiAgent — built new for that one call, never pooled or reused — starts a live voice conversation over Amazon Bedrock Nova 2 Sonic.
  • It guides them through an eleven-screen intake: history, medications, allergies, and a bounded 7–10 question, model-written symptom loop, one question at a time. Every screen transition is gated in code — the model cannot advance past a required field it hasn't collected.
  • Mid-call, the patient can reschedule, cancel, or upload a supporting document (like a lab report) without breaking the flow.
  • It never diagnoses, triages, or names a condition — it only listens and organizes. A standing notice states this at the start of every call.
  • The instant the call ends, a second, separate agent — Claude Sonnet 4.6 on Bedrock, with structured output — turns the raw transcript into a clean clinical-style report, including a type-only description of any uploaded document ("this appears to be an X-ray," never a finding).
  • The PDF lands in S3 and gets attached straight to the doctor's calendar event, alongside the appointment itself.

Healthcare is the first domain. The pipeline is built to generalize: the same engine, pointed at a sales discovery call or a financial advisory meeting, is a configuration change, not a rebuild.

How we built it

  • Strands Agents SDK end to end, on Amazon Bedrock — Nova 2 Sonic for the live bidirectional voice agent, Claude Sonnet 4.6 for the async summarization agent.
  • FastAPI edge, split by concern rather than by resource: booking & scheduling, session & tickets, doctor OAuth, reports & uploads, and a WebSocket gateway (/ws/intake/{session_id}) deliberately mounted outside /api/v1, because a live socket isn't a REST resource and pretending otherwise just hides the real architectural seam.
  • React 19 + Vite frontend: a booking wizard, a reschedule flow, a report lookup, and a guided call UI that surfaces one live question at a time.
  • Every action the live agent can take — navigating screens, recording an answer, rescheduling, requesting an upload, ending the call — is a narrow, individually gated Strands tool. Nothing is trusted to "the model behaving itself": navigate_to_screen literally refuses to advance past an unanswered required field, and end_session is backed by an 8-second force-end watchdog in case the model says goodbye and keeps talking anyway.
  • Security as one shape, not three: the intake link, the session cookie, and the WebSocket ticket are all HMAC-SHA256, from a single signer, with each token's purpose baked directly into what gets signed — so an intake token can never be replayed as a session token or a ws-ticket.
  • MongoDB Atlas for session/state data, Amazon S3 for reports, uploads, and recordings — all through presigned URLs, never proxied through our own API — and Google Calendar, via a service account joined by per-doctor OAuth for anything scoped to that physician.

Challenges we ran into

  • Choosing the right voice model, for the right reasons. Our first design targeted Gemini Live. We moved to Nova 2 Sonic specifically because Strands' own documentation demonstrated transcript capture working against it in real examples, and because Nova Sonic runs natively on Bedrock — directly benefiting from the hackathon's own AgentCore scoring incentive. We also had to catch that the legacy v1 Nova Sonic model was already marked EOL days after our own submission deadline, and pin v2 instead.
  • A real, live-only bug. Bedrock's own ValidationException message for a rejected image contained non-ASCII Unicode. Our structured logger's console renderer choked on it under Windows' default cp1252 encoding, silently turning a caught, handled error into an unhandled crash — but only against the real provider, never against our test fixtures. It's a good reminder that "works against mocks" and "works against production" are different claims.
  • A hard product boundary, decided the hard way. We originally had a safety-escalation path that tried to gauge urgency mid-call. We removed it entirely once we recognized that assessing urgency is triage, and this service isn't licensed to perform triage — replacing it with a standing notice stated once at the start of every call, rather than a judgment call left to the model.
  • Our own documentation drifted from our own code. Auditing our architecture diagram against the running system, we found two "in-call rebooking" tools we'd documented were dead code — the live agent had moved on to differently-named tools weeks earlier. The diagram wasn't wrong when we wrote it; it just went stale the way documentation always does. We now treat it as a claim to re-verify, not a memory to trust.

Accomplishments that we're proud of

  • A live, working, end-to-end voice agent — not a chat demo — that books, converses, gates its own behavior in code, and hands off a structured report, with zero shared state between concurrent calls.
  • A security model where three different token types share one signing scheme instead of three different ways to get authorization subtly wrong.
  • A product boundary (organize, never diagnose) that we were willing to cut a whole feature to enforce, rather than soften.
  • Recognizing, partway through the build, that the thing we'd built was bigger than the use case we'd built it for — and repositioning around that honestly instead of overselling healthcare as the whole story.

What we learned

That reliability in an agentic system comes from what you refuse to leave to the prompt — gated tools, bounded loops, watchdog timers, one-instance-per-conversation isolation — far more than from how well the model was asked to behave. And that the moment you add a new capability (a multimodal call, a new tool, a new integration), you should assume it will break in a way your existing tests can't see, because it usually does.

What's next for Context as a Service

Healthcare is the proof, not the ceiling. Next is pointing the same engine at a sales discovery call and a financial advisory meeting — same completeness gating, same gaps-over-guesses discipline, different domain configuration. Longer term, we want CaaS to be callable as an API: the thing any calendar or booking tool reaches for right before any two people meet for the first time.

Built With

Share this project:

Updates

Submission history