Headline FieldFuse: A deterministic ML data-compiler utilizing calibrated logistic regression and phonetic entity resolution to eliminate duplicate reporting for frontline healthcare workers.

Inspiration Frontline health workers in developing regions are drowning in administrative burden. A single antenatal care (ANC) visit can require a worker to enter the exact same data into three or four parallel systems (paper registers, local apps, and national databases). Recent studies document this duplicate reporting burden, leading to burnout, data silos, and clerical errors. We asked ourselves: How can we use Machine Learning not just as a black-box predictor, but as a verifiable data-compiler that safely bridges unstructured evidence into multiple structured systems?

What it does FieldFuse is a Python-based backend architecture (fieldfuse_core) and an offline-first PWA frontend that acts as an evidence-preserving compiler for health data.

Unstructured Capture: It takes unstructured inputs (like Live ASR voice transcripts or quick form hints) and extracts candidate facts. ML Entity Resolution: It uses a calibrated Machine Learning model (Scikit-Learn Logistic Regression with CalibratedClassifierCV) and fuzzy phonetic matching (rapidfuzz) to link these unstructured inputs to existing patient registers. Deterministic Safety (Safe Reject): FieldFuse never guesses. If our ML detects a conflict (e.g., a patient claims their last visit was August 1st, but the resolved register says July 15th), it triggers a Safe Reject—a hard block requiring human-in-the-loop verification. Compile Everywhere: Once the human verifies the facts, the engine compiles the canonical event into multiple backend formats simultaneously (e.g., ABDM FHIR R4, ANMOL, U-WIN demo replicas). How we built it The core technical architecture was built in Python using a highly deterministic pipeline design:

The Entity Resolver Engine: We trained an ML model using scikit-learn to perform fuzzy probabilistic matching. We engineered custom features including token_set_ratio, custom phonetic sound keys (normalizing vowels and hard consonants), and Date-Of-Birth swap-error detectors (day/month transposition). Probability Calibration: Because healthcare requires strict confidence thresholds, we wrapped our Logistic Regression model in a CalibratedClassifierCV. We optimize for the brier_score_loss to ensure our predicted probabilities reflect true match confidence. If the confidence falls below our strict threshold, the engine safely abstains. Deterministic Verification Engine: We wrote a strict Python rule engine (conflicts.py) that generates ResolutionRequests. An unresolved HARD_BLOCK physically prevents the system from compiling the final JSON adapters. 3D Visualized Provenance: On the frontend, we built an interactive WebGL/Three.js architecture. The canvas visualizes the ML pipeline: an orbiting "evidence" node physically crashes into the canonical core if a conflict is detected, visually proving the data provenance to the end user. Challenges we ran into Building an ML architecture for healthcare is terrifying because false positives (linking medical records to the wrong person) can have disastrous clinical consequences. We struggled initially with out-of-the-box ML models producing overconfident false positives. To solve this, we had to implement rigorous Brier score probability calibration and hardcoded phonetic logic specific to Indian frontline names. Another major challenge was building a system where the ML output isn't immediately accepted; designing a deterministic "compiler" engine that sandboxes ML predictions until they pass conflict evaluation required a complete rethink of standard ML pipelines.

Accomplishments that we're proud of We are incredibly proud of our Entity Resolution Architecture. By combining rapidfuzz string metrics with date transposition logic and a calibrated classifier, we achieved stellar Precision-Recall metrics on our synthetic validation split. Furthermore, we achieved 100% deterministic replay on our test suite. We can take an entire offline session of voice transcripts, run them through the ML resolver, and guarantee that 10/10 compiles produce the exact same final FHIR JSON digests.

What we learned We learned that in healthcare, ML shouldn't be a black box that makes final decisions. Instead, ML is best used as an evidence gatherer. We learned how to build a highly constrained "trust architecture" where the core Python engine treats ML outputs as candidates rather than facts. We also mastered integrating Scikit-Learn probability calibration directly into a deterministic data-pipeline.

What's next for FieldFuse We plan to expand our local subset validations for FHIR exports, introduce more sophisticated edge-LLM models for offline medical document scanning, and expand our ML feature engineering to handle complex multilingual phonetic spelling variations across different Indian dialects.

Share this project:

Updates

Submission history