Inspiration

In India, getting to the right care is often not just a medical problem. It is an information problem.

Families and care coordinators may know that help exists somewhere nearby, but in a high-stakes moment they still have to answer difficult questions quickly: Which facility can actually handle this case? How far away is it? What evidence supports that recommendation? What information is missing? What should be verified before sending the patient?

Referral Copilot was inspired by this gap between “there are hospitals on a map” and “this patient has a realistic path to care right now.” We wanted to build an AI-powered referral navigator for underserved and resource-constrained communities in India, where delays, fragmented data, language barriers, and uncertainty can turn an urgent need into a dangerous waiting game.

Our goal was not to replace clinicians or coordinators. It was to give them a transparent, evidence-backed copilot when every hour matters.

What It Does

Referral Copilot turns a plain-language request like:

dialysis near Jaipur

or

emergency surgery near Patna

into an explainable referral workflow.

The app:

  • Parses the care need, urgency, and location from natural language.
  • Supports free-text search, pin-code search, and device-location search.
  • Resolves geography using facility addresses, pincode data, and optional LLM-powered web resolution when local data is incomplete.
  • Ranks nearby facilities on a transparent 0-10 suitability score.
  • Shows an interactive Leaflet map with numbered hospital markers, route lines, popups, and click-to-inspect facility details.
  • Displays matching evidence from facility records such as specialties, procedures, equipment, capability, and descriptions.
  • Flags missing or suspicious evidence such as absent phone numbers, missing websites, unclear emergency readiness, sparse specialties, missing source URLs, or invalid coordinates.
  • Enriches candidates with LLM-assisted web discovery, including public rating/review signals, official hospital websites, and source links when confidently available.
  • Lets coordinators save facilities to a shortlist, download the shortlist as CSV, and compare options.
  • Provides a shortlist copilot that can answer questions, verify public details, explain tradeoffs, and generate coordinator next steps.
  • Supports multilingual UI and result display for Indian languages, while preserving hospital names, phone numbers, URLs, pin codes, coordinates, and scores.
  • Includes dark/light mode and placeholder Call/Schedule buttons for future referral workflow integration.

The app is designed as referral decision support, not medical advice. It helps coordinators move faster while still showing the evidence and uncertainty behind each recommendation.

How We Built It

We built Referral Copilot as a Databricks App using Python, Dash, Dash Leaflet, pandas, Plotly, the Databricks SQL Connector, and OpenAI.

Databricks acts as the trusted data layer. The app reads from Unity Catalog tables in the databricks_virtue_foundation_dataset_dais_2026 catalog, including facility data, India pincode geography, and NFHS-5 district health indicators.

OpenAI is used carefully as an assistive layer, not as the source of truth. It helps with:

  • Natural-language query parsing.
  • Spelling correction for locations.
  • Translating non-English user inputs into English matching terms behind the scenes.
  • Multilingual UI/result display.
  • Web-assisted enrichment for official websites, public review signals, and source verification.
  • Shortlist copilot answers with optional web search when fresh context is needed.

The ranking engine itself is deterministic and evidence-grounded.

Scoring Logic

Each facility receives a base score on a 0-10 scale. The raw score combines facility evidence, distance, facility type, urgency fit, and data-quality penalties:

$$ R = E + D + F + U - M $$

The base score is then scaled:

$$ S_{base} = \mathrm{clamp}_{0}^{10}\left(\frac{R}{12}\right) $$

Where:

  • (E) = evidence score from matched facility fields.
  • (D) = distance score.
  • (F) = facility type bonus.
  • (U) = urgency fit bonus or penalty.
  • (M) = missing/suspicious evidence penalty.

Evidence fields have different weights:

Field Weight
Specialties 10
Procedure 9
Equipment 8
Capability 8
Description 5
Name 3
Facility type 3
Operator type 1

Distance is rewarded more strongly for nearby facilities:

$$ D = \begin{cases} 70, & \text{if distance} \leq 5 \text{ km} \ \max\left(5,\ 70 \times \left(1 - \frac{\text{distance}}{\text{radius}}\right)\right), & \text{otherwise} \end{cases} $$

For example, suppose a user searches for dialysis near Jaipur.

A hospital 8 km away with dialysis evidence in specialties and equipment might receive:

$$ E = 10 + 8 = 18 $$

If the search radius is 25 km:

$$ D = 70 \times \left(1 - \frac{8}{25}\right) = 47.6 $$

If it is a hospital:

$$ F = 8 $$

If it has one missing signal, such as no official website:

$$ M = 2.5 $$

For a routine dialysis search:

$$ U = 0 $$

So:

$$ R = 18 + 47.6 + 8 + 0 - 2.5 = 71.1 $$

$$ S_{base} = \frac{71.1}{12} = 5.9 $$

The app may then apply a small public-review adjustment from LLM-assisted web enrichment:

$$ S_{final} = \mathrm{clamp}{0}^{10}(S{base} + P) $$

Where (P) is a bounded public rating/review modifier. Public signals are shown separately because reputation is useful context, but not clinical evidence.

For example, if a hospital has a strong public rating signal and enough review confidence:

$$ P = +0.7 $$

Then:

$$ S_{final} = 5.9 + 0.7 = 6.6 $$

This makes the score understandable: clinical/facility evidence and distance drive the ranking, while public web signals can gently adjust the final ordering without replacing the source data.

Challenges We Faced

Healthcare data is messy. Some facility records had missing websites, missing phone numbers, sparse specialties, directory-style text, malformed coordinates, or source fields that needed cleanup. Some locations were easy to resolve with pincode data, while others required fallback matching from facility addresses or web-assisted resolution.

The biggest challenge was trust. In healthcare, a polished AI answer is not enough. The app needed to show why a facility was recommended, what evidence matched, what was missing, and what a coordinator should verify before referral.

We also had to balance live web context with reliable dataset evidence. Public ratings, official websites, and review snippets are helpful, but they should not be confused with clinical capability. That is why Referral Copilot separates public signals from facility-record evidence and always surfaces uncertainty.

Another challenge was multilingual access. India is not one-language-first. We added a global language selector so coordinators can use the app in regional languages, while preserving critical identifiers like hospital names, phone numbers, links, and pin codes.

What We Learned

We learned that AI is most valuable in healthcare when it is grounded, humble, and operational. The win is not just answering a question. It is helping someone compare options, identify evidence gaps, make the next call, and move a patient through a fragmented system faster.

We also learned that referral navigation is both a data integration problem and a human trust problem. Maps, scores, evidence snippets, missing-data warnings, public web context, and chat become much more powerful when they work together as one workflow.

Accomplishments We Are Proud Of

  • Built a working Databricks App for AI-assisted referral navigation.
  • Created an evidence-ranked facility shortlist on a transparent 0-10 scale.
  • Added interactive Leaflet maps with numbered markers, route lines, and facility popups.
  • Built a shortlist copilot that can compare facilities, explain tradeoffs, verify details, and generate next steps.
  • Added LLM-assisted web enrichment for public signals, source links, and official websites.
  • Added multilingual app support for Indian languages.
  • Added light/dark mode and a more polished coordinator-focused UI.
  • Included safeguards for missing evidence, noisy records, invalid coordinates, API rate limits, and uncertain web results.
  • Kept deterministic facility evidence separate from LLM-generated context.

What's Next

Next, we would like to add real-time bed availability, ambulance availability, emergency department readiness, insurance and eligibility checks, voice input, WhatsApp/SMS handoff, and deeper integration with local public health networks.

We would also expand the dataset coverage nationally, incorporate more district/state-level health indicators, and track referral outcomes so the system can learn which recommendations actually led to care.

The long-term vision is a referral navigation layer for underserved communities: fast enough for emergencies, transparent enough for trust, and practical enough for coordinators working under pressure.

Built With

Share this project:

Updates

Submission history