Inspiration
India carries the world's largest tuberculosis burden — over 2 million new cases every year — and TB remains the country's #1 infectious killer. The first line of defense is not a hospital: it's community health workers like India's ASHA workers, who decide which of the coughing people they visit should be referred for confirmatory testing.
But at the frontline there is no lab, no Xpert machine, no radiologist, no doctor. Self-reported cough history is unreliable, and symptom checklists alone fall short of the WHO's target product profile for a TB triage test.
AI cough analysis looked like the answer — until you read the external validations. The largest effort to date, the CODA TB DREAM Challenge (733,756 cough sounds, 2,143 patients, 7 countries, Nature Scientific Data 2024), saw models that performed well in-country collapse when tested across borders: AUC dropped from 0.69 to 0.48–0.62 in Peru external validation. Cough AI works — just not everywhere, and not for the people who need it most.
That gap became SvaraTB. ("Svara" = Sanskrit for "sound", honoring the IISc Coswara dataset that pioneered Indian cough research.)
What it does — the proposed solution
SvaraTB turns a health worker's ordinary smartphone into a TB triage companion. It answers exactly one question: who should be referred for confirmatory testing first?
Three complementary signals are fused on-device:
- Passive cough frequency (48h) — a lightweight on-device cough detector counts coughs per hour in the background. Proven signal (Uganda smartphone studies, medRxiv 2025, AUC 0.69–0.76) but too weak alone — we use it as one modality, not the answer.
- One guided cough — the worker prompts a single deep cough and records 3 seconds. Speech foundation models (WavLM / Audio Spectrogram Transformer) extract acoustic embeddings — the approach a 2025 Zambia study (arXiv 2509.09746) showed can reach WHO triage targets.
- Structured clinical inputs — ~12 dimensions of symptoms and demographics the worker already collects.
The output is not a probability to interpret — it's a calibrated three-tier recommendation: REFER for confirmatory testing / LOW RISK / UNCERTAIN, aligned to the WHO target of 90% sensitivity. "Uncertain" is shown honestly instead of guessed.
Two properties make this different from every cough-AI-before-it:
- Audio never leaves the phone. No upload, no server, no consent nightmare.
- The model adapts to each region without seeing regional data. Phones fine-tune locally and share only model gradients — federated domain adaptation — attacking the cross-border generalization collapse head-on.
Technology component — how it works
[Passive listening 48h] lightweight ONNX cough detector → coughs/hour
│
[Guided cough, 3s] → DC removal / 80Hz high-pass / silence trim /
peak-norm / resample 16kHz
│
▼
[On-device embedding] WavLM / AST (int8 ONNX) — runs on mid-range phones
│
▼
[Multi-modal fusion] embedding ⊕ frequency features ⊕ 12-dim clinical
→ lightweight classifier (XGBoost / small MLP)
│
▼
[Calibrated decision] Platt / temperature scaling → 3 tiers
aligned to WHO TPP (≥90% sensitivity)
│
▼
[Referral guidance] localized next-step instructions (Hindi, Bengali…)
│
▼
[Federated adaptation] local fine-tune → gradients only → server FedAvg
→ updated global weights. Audio never leaves device.
Why this is feasible — every piece already exists:
| Component | Evidence it works |
|---|---|
| Passive cough counting on phones | Uganda studies (Hyfe), medRxiv 2025 |
| Foundation-model cough embeddings reach WHO TPP | Zambia study, arXiv 2509.09746 |
| On-device inference | WavLM int8 ONNX runs on mid-range phones (Jaga prototype, AMD hackathon) |
| Training data | Public: CODA TB (733k coughs, 7 countries), Coswara (IISc, 2,635 people), COUGHVID (30k+), Virufy |
| Federated averaging | Standard, mature method (FedAvg variants) |
No new hardware. No new dataset collection to start. The science risk is concentrated exactly where our contribution is.
How we built it
This is an ideathon concept, so the deliverable is the design — but it is grounded in verified literature rather than speculation. We surveyed every serious prior attempt (CODA TB, the Zambia foundation-model study, Uganda passive-frequency work, the Jaga prototype) and found that each validated one component while leaving the others' weaknesses exposed. SvaraTB's contribution is the composition: fusing the three signals so each covers the others' blind spots, and adding federated domain adaptation where CODA TB's models demonstrably failed.
Challenges and limitations — honestly stated
- Cross-region generalization is an open research problem. We do not claim to have solved it; we propose federated domain adaptation as a concrete, testable path, plus a validation plan (train on CODA TB's 7-country split, hold out countries as external tests).
- Passive listening raises privacy questions. Mitigation: the detector stores event counts only, never audio; users see exactly what is counted.
- Specificity at 90% sensitivity is the known trade-off (Uganda showed 20–46% specificity at high sensitivity). Multi-modal fusion exists precisely to push that number up — and the three-tier output prevents over-referral from becoming alarm fatigue.
- It is not a diagnosis. SvaraTB does not replace sputum tests, Xpert, or clinical judgment; it decides who gets the next test. Confirmatory pathways stay unchanged.
Target users and potential impact
- Primary users: ASHA workers and community health volunteers in high-burden regions — the people who today make referral decisions with no tools at all.
- Beneficiaries: suspected TB patients in remote and low-income communities — earlier triage means earlier treatment, lower mortality, and less onward transmission.
- Scale: the same pipeline extends to pneumonia, asthma, and other respiratory disease surveillance — any condition where "who needs the scarce test first" is the real question.
What's next — implementation plan
- Phase 1 — core triage prototype: passive counter + guided cough + on-device fusion, trained on CODA TB, validated on Coswara (the Indian domain). Zero new data collection required.
- Phase 2 — federated adaptation: simulated multi-region clients (CODA TB's per-country splits are a natural testbed) to measure whether gradient-sharing closes the external-validation gap.
- Phase 3 — field partnership: pilot with frontline health programs (e.g., integration with India's National TB Elimination Programme), ethics review, and closed-loop calibration against Xpert results.
Accomplishments we're proud of
- A design that turns the field's biggest documented failure — cross-border collapse — into the core contribution rather than a footnote.
- Every technical claim is backed by a public dataset or a published result; nothing requires inventing new science to evaluate.
- Privacy is architectural, not promised: audio physically cannot leak because it never leaves the device.
What we learned
The literature showed us that "cough → TB" models already exist — what doesn't exist is a model that keeps working when it crosses a border, and a deployment model that rural health systems can actually accept. The hard problem wasn't classification accuracy; it was generalization, calibration, and trust.
AI-use disclosure
AI tools were used as a research and drafting aid in preparing this submission. The problem analysis, solution composition, technical reasoning, and limitations assessment are the author's own work.
Built With
- audio-signal-processing
- edge-computing
- federated-learning
- machine-learning
- mobile-health
- onnx
- privacy
- speech-models
- tuberculosis
- xgboost
Log in or sign up for Devpost to join the conversation.