Inspiration
I'm a statistician at a teaching hospital in Nigeria, and I've spent years around maternal health data that shows the same pattern over and over: Nigeria's own national health survey found that 63% of women attend antenatal care, but only 46% have a skilled attendant at delivery. An earlier DHS survey found the gap even wider — 74% ANC attendance against 41% facility delivery. That gap has a name, the ANC Paradox, and it's been documented since at least 2018. The problem was never detection — the risk factors are well studied. It was that no one was watching at population scale, in the background, without someone manually reviewing every record.
What it does
ANC Fleet is an async agent pipeline, not a chatbot. It runs in the background against a batch of ANC records: a Scanner Agent pulls the batch, a Risk Agent scores every patient against risk weights grounded in published Nigeria DHS findings (education, wealth quintile, residence, parity, maternal age, region, ANC completion), and an Action Agent schedules outreach for anyone flagged — skipping anyone already contacted in a previous run, so the pipeline is genuinely stateful across weeks rather than a one-shot demo. Gemini then generates a plain-language brief a health worker could read in 20 seconds, turning a batch of risk scores into something actionable.
Every stage — scan, score, decide, act — writes a timestamped record to Firestore. That audit trail was a deliberate design choice, not an afterthought: it's what lets anyone verify the pipeline actually ran, rather than trusting a diagram that claims it does.
How I built it
The scoring model is deterministic, transparent Python — not a black box, and not an LLM guessing at risk. Every weight traces back to a cited DHS finding: for example, adequate ANC attendance ranges from 34.6% (no education) to 92.3% (higher education), and from 30.7% (poorest wealth quintile) to 89.2% (richest). I translated those documented gradients into a weighted scoring function, then validated it against the DHS's own documented best-case and worst-case demographic profiles.
Getting the calibration right took real iteration. My first pass flagged 83% of a test population as needing outreach — far too aggressive to read as a targeted intervention tool rather than a blanket alert. I recalibrated against the real national non-facility-delivery rate (~54-59%, since NDHS reports 41-46% facility delivery) and added an automated regression test so the model can't silently drift back toward over-flagging as the code changes. That test runs alongside four others verifying the scoring logic end to end.
Gemini's role is deliberately scoped to what it's good at: turning structured results into plain language, not making the underlying risk decision. That split — deterministic, auditable scoring plus LLM-generated communication — was a conscious choice for a healthcare context where "why was this patient flagged" needs a traceable answer.
The pipeline is built with Google's Agent Development Kit (ADK) for agent structure, Gemini for the outreach brief, and Firestore for the audit trail and cross-run patient memory.
Challenges I ran into
The biggest challenge wasn't code — it was infrastructure access. I originally built this against Vertex AI and planned to deploy on Cloud Run, but hit a regional billing-verification wall specific to Nigerian accounts that I couldn't resolve within the hackathon window, even after trying multiple GCP projects and requesting help from a collaborator. Rather than stall, I switched the Gemini integration to the direct Gemini API (Google AI Studio key), which doesn't require the same billing verification, and deployed the Flask service to Render instead of Cloud Run. Firestore — real Google Cloud infrastructure — and Gemini both run live in this submission; only the compute host differs from my original plan, and that difference is documented transparently in the repo's README rather than hidden.
I also learned a smaller but important lesson mid-build: my first deployment had a bug where Render's routine health-check pings were hitting the same endpoint as the real pipeline trigger, silently burning through my Gemini API quota. Fixing it meant splitting the health check (GET) from the actual trigger (POST) — a small change, but the kind of thing that only shows up once something is actually deployed and being monitored, not just running locally.
What I learned
That transparent, cited logic is more defensible under scrutiny than a bigger model would have been — when a judge asks "why was this patient flagged," I have a specific, sourced answer, not a black box. And that infrastructure constraints are real engineering problems worth solving and documenting, not just embarrassing gaps to paper over.
What's next for ANC Fleet
The immediate next step is swapping the synthetic, NDHS-shaped population for real facility data via CliniqBridge, a FHIR R4 MCP server I built and already run in production — that swap is a one-line configuration change, not a rewrite, once institutional ethics approval for real patient data is granted. Longer term: wiring in Cloud Scheduler for true unattended async triggering (currently manually triggered via POST) once Cloud Run access is resolved, and piloting real SMS/call outreach through the existing SentinelCall pipeline, currently running in dry-run mode.
Built With
- agents
- async
- fhir
- firestore
- flask
- gemini
- google-adk
- google-ai-studio
- google-cloud
- healthcare
- python
- render
Log in or sign up for Devpost to join the conversation.