Inspiration

Nearly 1 in 4 Americans is now a family caregiver. They spend an average of 27 hours per week coordinating care for aging parents — chasing refills, rescheduling appointments, managing deliveries, and fielding pharmacy calls. Existing tools are fragmented: a medication app, a calendar app, a delivery app, a messaging app. None of them talk to each other. The adult child becomes the integration layer. CareBridge is built for them.

What it does

CareBridge is a production-grade multi-agent system built with the Strands Agents SDK. It runs in the background, monitoring medication adherence, pharmacy refills, appointments, and deliveries. When an event requires action, the Supervisor Agent — powered by a real LLM (provider-agnostic: Groq in the current deployment, plus Gemini, Bedrock, Anthropic, OpenAI, or local Ollama) — reasons about the event, routes to the correct specialist agent, and executes autonomously. The caregiver only sees what needs a human decision.

Five agents work together: • Supervisor Agent — orchestrates, classifies actions, maintains an immutable audit trail • Medication Agent — monitors adherence, orders refills, detects deviation patterns • Appointment Agent — manages provider calendars, coordinates transportation, sends prep checklists • Logistics Agent — coordinates grocery and pharmacy deliveries, monitors failures • Communication Agent — generates proactive family alerts, handles two-way queries

Every action is logged to a tamper-proof audit trail before execution. Deterministic escalation logic — not the LLM — decides what requires human approval.

How we built it

• Strands Agents SDK — agents-as-tools pattern with a real LLM-driven Supervisor • LLM provider-agnostic layer — one env var switches between Groq, ♊ no, Bedrock, Anthropic, OpenAI, or Ollama • FastAPI backend — JWT auth, RBAC, rate limiting, structured JSON logging • Firebase Auth — Google sign-in verified server-side via firebase-admin, backend issues business-claim JWTs • Next.js 15 dashboard — real-time care status, alert feed, approval queue, audit viewer, live chat with the Supervisor • Four MCP servers (Qoder Agent SDK, in-process) — pharmacy, messaging, delivery, calendar. Standalone modules with registered tools; the Strands Supervisor currently calls the underlying Python tool functions directly • Immutable audit trail — SQLite with triggers blocking UPDATE and DELETE • 177 automated tests — unit + integration + API + LLM routing + audit immutability • Dockerized — multi-stage build, non-root runtime, healthcheck

Challenges we ran into

The hardest part was making the Supervisor a real agent rather than a router. Early versions used deterministic if/elif routing — functional, but not an agent. We rewrote the Supervisor to perform a real inference turn: the LLM sees the four specialist agents as tools and decides which to invoke.

Proving the LLM was genuinely in the decision path became a design constraint. We wrote a permanent AST-based test that parses supervisor_agent.py and fails the build if anyone reintroduces if event_type == routing. That test runs on every push.

The second challenge was deployment. Firebase service account credentials cannot be committed, and long JSON strings get mangled when pasted into cloud environment variables. We added a base64-encoded credential path (FIREBASE_CREDENTIALS_B64) that survives any paste, alongside the existing file-path and JSON-string fallbacks.

The third was honest scope. Pharmacy and delivery APIs require commercial contracts — they cannot be integrated in a hackathon window. Rather than pretend, we documented exactly what is real and what is simulated. Every mocked adapter is swappable without touching agent code.

Accomplishments that we're proud of

  1. A real LLM-driven Supervisor — not a router. The Supervisor performs a real inference turn on every event. The LLM sees the four specialist agents as tools and decides which to invoke. We enforce this with a permanent AST-based test that fails the build if anyone reintroduces deterministic event_type routing.

  2. Provider-agnostic LLM layer — bring your own model. One environment variable switches between Groq, Gemini, Bedrock, Anthropic, OpenAI, and local Ollama. CareBridge is not locked to any single vendor. Enterprise-grade flexibility was designed in from the start, not bolted on.

  3. Tamper-proof audit trail. Every agent action is written to an immutable SQLite audit log BEFORE execution — enforced by triggers that block UPDATE and DELETE. For a care coordination agent, observability is not a feature; it is the trust layer.

  4. Deterministic escalation logic — safety over cleverness. The LLM routes tasks. Python decides what requires human approval. This separation is enforced by architecture: LLMs cannot classify their own actions as autonomous. Escalation is testable, auditable, and injection-proof.

  5. Five agents, one production system. Supervisor + Medication + Appointment + Logistics + Communication agents, each with a scoped tool surface, orchestrated through the Strands agents-as-tools pattern. Four standalone MCP servers (pharmacy, messaging, delivery, calendar) exist as in-process modules built with the Qoder Agent SDK; native MCP wiring into the Strands runtime is on the roadmap.

  6. 177 automated tests passing — zero regressions. Unit tests, integration tests, API tests, LLM routing tests, audit immutability tests, Firebase auth tests, CRUD tests. Every push must pass. The build is not a demo; it is a system with a test suite.

  7. Production-grade backend, not a prototype. FastAPI with JWT auth, refresh tokens, RBAC, rate limiting, request-ID middleware, structured JSON logging, global exception handling, request size limits, health/readiness/version endpoints, and an OpenAPI spec. Dockerized with a multi-stage build, non-root runtime user, and healthcheck.

  8. Realistic constraints documented honestly. Pharmacy and delivery APIs are mocked because production credentials require signed contracts — an industry constraint, not a build limitation. The README contains an explicit "What Is Real vs Simulated" table. Every mocked adapter is swappable without touching agent code.

  9. A dashboard a sibling would actually use. Calm, mint-toned, data-dense but readable. Every metric is actionable; every alert is tied to an approval or notification. Built to reduce caregiver anxiety, not add to it.

  10. Shipped end-to-end in under 72 hours. From empty folder to 177 passing tests, a working dashboard, a production backend, four MCP servers, Firebase auth, and a provider-agnostic LLM layer — in a single weekend. The system is deployable today.

    What we learned

  11. Deterministic escalation is non-negotiable for care coordination. The LLM routes, but Python decides what requires human approval.

  12. Provider-agnosticism should be a day-one design decision, not a retrofit. Bolting it on later cost significantly more.

  13. The audit trail is the product. For a care coordination agent, observability is not a feature — it is the trust layer.

  14. Honesty about scope is a strength. Disclosing what is simulated (pharmacy, delivery) gives credibility to what is real (agent orchestration, safety policy, audit trail).

    What's next for CareBridge — The AI Coordinator for Aging Parents• Real pharmacy API • Native MCP wiring into the Strands runtime (currently the Supervisor

    calls the underlying tool functions directly) • Real Twilio SMS integration (trial credentials ready, single-file swap) • Google Calendar OAuth with the calendar.app.created scope (avoids Google verification delay) • Real pharmacy API integration (requires commercial contracts — Surescript or Walgreens API) • Real delivery API integration (requires commercial contracts — Instacart or DoorDash API) • HIPAA-ready deployment on Amazon Bedrock AgentCore • Multi-recipient support for caregivers managing more than one parent

Built With

Share this project:

Updates

Submission history