Inspiration

If you've ever called a dental or cosmetic clinic and gotten voicemail, or sat on hold while the one person at the front desk juggles three other calls, you already know the problem. I've built voice AI for clinics before as a freelancer, and the pattern is always the same: the front desk isn't understaffed because anyone's bad at their job --- it's understaffed because the phone never stops ringing, and most of what it rings about is repetitive. Confirm an appointment. Answer "do you take my insurance." Chase a patient who might not show up.

When AWS announced Agents for Humans with the Strands Agents SDK, I wanted to build something that actually runs the front desk in the background, not another chatbot that just answers questions.

What it does

ClinicPilot is a voice-first AI front desk agent. A patient talks to it through a browser microphone --- no phone line needed --- and it checks availability, books or reschedules appointments, and answers clinic-specific FAQs from the clinic's own knowledge base.

Separately, it runs autonomously: a daily background job scans upcoming appointments and decides whether to send a reminder or flag a patient with a no-show history for staff. It never reschedules on a patient's behalf --- that's a decision a human should make --- so the autonomy has a hard boundary by design.

It's built multi-tenant from the data layer up. The demo proves this with two seeded clinics, one dental and one cosmetic, served by the same deployed agent with different hours, services, and FAQ content --- not two separate deployments.

There's also a staff dashboard: the live appointment calendar, an escalation queue with email alerts, and schedule settings the voice agent books against.

How I built it

The core is an orchestrator agent plus three specialist sub-agents --- Scheduling, FAQ, and Escalation --- wired together with Strands' Agent-as-Tool pattern. Each sub-agent is a real Strands agent exposed to the orchestrator as a callable tool, which gave genuine multi-agent orchestration without the deployment overhead of running each one as a separate service.

Voice runs on Amazon Nova Sonic paired with Strands' BidiAgent for bidirectional audio streaming --- a real speech-to-speech conversation instead of a stitched transcribe-then-synthesize pipeline. I started from AWS's own sample implementation as a scaffold rather than building the streaming plumbing from scratch.

Tenant isolation is enforced at the tool boundary: every tool function requires clinic_id, and every DynamoDB query and Bedrock Knowledge Base lookup is scoped to it.

The background automation (EventBridge Scheduler + Lambda) calls the exact same backend/tools/ functions the live voice agent uses, so there's one source of truth for business logic, not two.

Everything is deployed on Amazon Bedrock AgentCore (Runtime + Memory), with the full stack --- DynamoDB, Lambda, EventBridge, API Gateway, Cognito, SES, S3/CloudFront --- defined in AWS CDK (Python).

Both the Python and TypeScript sides were built test-first: the tool layer, the transcript state machine, and the infra parity are all covered by unit suites that run without touching AWS.

Challenges I ran into

1. A floating dependency silently broke the deployed runtime

An unrelated redeploy rebuilt the agent container from scratch, and the unpinned strands-agents requirement resolved to a version released that same week, one that renamed the Nova Sonic model class behind a lazy import.

Local code kept working because my virtual environment had the older wheel. The deployed runtime failed every call with an ImportError that surfaced to the browser only as a closed WebSocket.

The fix was pinning the requirement and learning the lesson: with a floating dependency floor, every container rebuild can become a silent upgrade.

2. The same model ID needed two different IAM ARNs

Bookings failed with AccessDenied even though the role looked correct.

A live probe captured the sub-agent's actual tool_result: Strands invokes the model against a region-less foundation-model ARN, while Nova Sonic's bidirectional stream authorizes against a region-scoped one.

So the same model ID requires both ARN shapes to be granted, and a policy with only one can never match the other.

I confirmed it with iam simulate-principal-policy before and after the fix.

3. The transcript state machine didn't match the wire

Probing the live event stream showed the real contract differs from what any sample assumed: patient speech arrives as a single final event containing the complete utterance (no interims), while assistant speech arrives only as sentence deltas that never finalize.

The first reducer duplicated every patient word and vanished agent replies mid-conversation.

I rewrote it test-first against the captured event shapes --- red tests on the old code, green on the new.

4. Browser auth for AgentCore has an invisible ceiling

The patient voice UI needs anonymous AWS credentials, but Cognito's enhanced auth flow applies a scope-down session policy whose service allow-list excludes Bedrock AgentCore entirely --- no role policy can override it.

The fix was the classic basic auth flow (GetOpenIdToken + STS AssumeRoleWithWebIdentity), which nothing in the error message points you toward.

Accomplishments I'm proud of

  • A working end-to-end voice conversation, fully deployed on AgentCore
  • The same deployed agent correctly serving two different clinics
  • A background agent that takes real autonomous action and only escalates to a human when it should, and never acts where a human should decide
  • A tenant-isolation boundary enforced at the code level, not just assumed

What I learned

The hardest part of building "an AI that answers the phone" was never the phone --- it was knowing when the agent should act and when it should stop and hand off to a person.

I also learned that starting from a reference implementation for the riskiest, most experimental part of the stack --- voice streaming --- is what made a solo build of this scope possible at all.

What's next

  • Real telephony via Amazon Connect
  • Self-serve clinic onboarding
  • Expanding the autonomous background rules beyond reminders and escalations

Built With

Share this project:

Updates

Submission history