MirAI

A Multi-Stage Machine Learning Cascade for Liquid Biopsy-Based Alzheimer's Disease Triage
Explore Live Portal » · API Docs » · Interactive Flowcharts »

Java 21 Spring Boot React 19 Python 3.12 FastAPI GCP Cloud Run


Table of Contents

  1. About The Project
  2. Machine Learning Rigor & Methodology
  3. System Architecture
  4. Project Structure
  5. Getting Started
  6. API Reference
  7. Evaluation & Results Summary
  8. Roadmap
  9. License & Disclaimer
  10. Acknowledgements

About The Project

Clinical Motivation & The Problem

Alzheimer’s Disease (AD) accounts for 60–70% of dementia cases worldwide and is projected to reach 152 million cases by 2050. While disease-modifying therapies are emerging, clinical triage remains a critical bottleneck.

Under the current standard of care, confirming early-stage AD pathology requires invasive lumbar punctures (CSF assays) or high-cost Amyloid Positron Emission Tomography (PET) scans—costing an estimated $45,531 per verified case. This creates long specialist waitlists and excludes resource-constrained populations from timely intervention.

The 3-Stage Clinical Cascade

MirAI introduces a machine learning-driven clinical gatekeeper trained on the Alzheimer’s Disease Neuroimaging Initiative (ADNI) cohort. It implements a progressive 3-stage risk escalation architecture that requires additional, higher-cost biomarker testing only when mathematically and clinically justified:

flowchart LR
    A["Stage 1: Demographic Screen\n(Age, Gender, Education)"] -->|Elevated Risk| B["Stage 2: Genetic Stratification\n(APOE-ε4 Allele Count)"]
    B -->|Elevated Risk| C["Stage 3: Plasma Liquid Biopsy\n(p-tau217/Aβ42, NfL, GFAP)"]
    C -->|High Calibrated Risk| D["Accelerated Referral\n(Structural MRI & Amyloid PET)"]
    A -->|Low Risk| E["Routine Monitoring"]
    B -->|Low Risk| E
    C -->|Low Risk| E
  1. Stage 1 (Primary Care Demographic Screen): Zero-cost baseline assessment using demographic factors (AGE, PTGENDER, PTEDUCAT).
  2. Stage 2 (Genetic Augmentation): Incorporates one-time APOE4 genotyping to stratify genetic predisposition.
  3. Stage 3 (Plasma Liquid Biopsy Panel): Ingests blood-based biomarkers (pT217_AB42_F, AB42_AB40_F, NfL_Q, GFAP_Q) via a calibrated gradient-boosted engine, detecting active molecular pathology before irreversible neurodegeneration occurs.

Built With

Component Technology Version Description
ML Inference Engine Python / FastAPI 3.12 / 0.115 High-performance asynchronous REST microservice
ML Models XGBoost / scikit-learn 2.1 / 1.8 Gradient boosting with native NaN routing & Isotonic Calibration
Model Explainability SHAP 0.50 Fold-aggregated TreeSHAP feature attribution
Enterprise Backend Java / Spring Boot 21 / 3.4.1 JWT authentication, REST gateway, assessment history
Frontend UI React / Vite 19 / 6.0 Responsive Single Page Application with Chart.js & Glassmorphism
Database MySQL / GCP Cloud SQL 8.0 Relational storage for users, predictions, and audit logs
Cloud Infrastructure Google Cloud Run Fully Managed Serverless containerized deployment across microservices

Machine Learning Rigor & Methodology

MirAI adheres to rigorous medical AI translational standards:

Leakage Prevention

  • Patient-Level Group Isolation: Enforced strict GroupShuffleSplit on unique Patient Identifiers (RID). Repeated observations from the same subject never cross train/test boundaries.
  • Purged Clinical Target Leakage: Diagnostic instruments (e.g., MMSE, CDR-SB, FAQ) used to define the diagnostic label (DX) were explicitly excluded from features to ensure honest pre-symptomatic triage.
  • Modality Integrity: Invasive CSF biomarkers were eliminated from Stage 3; only non-invasive plasma assays are used.

Missing-Not-At-Random (MNAR) Handling

In real-world cohorts, blood biomarker missingness reflects clinical decision patterns rather than random omissions. MirAI avoids synthetic imputation and instead leverages XGBoost's native NaN routing augmented with explicit missingness indicator features (_missing).

Calibration & Explainability

  • Isotonic Calibration: Calibrated probability outputs via CalibratedClassifierCV to minimize Expected Calibration Error (ECE ≈ 0.12).
  • Cross-Fold SHAP Aggregation: Model explainability is computed inside cross-validation partitions and aggregated, identifying pT217_AB42_F and NfL_Q as the dominant biological predictors of progression risk.

System Architecture

┌──────────────────────────────────────────────────────────────────────────────────┐
│                            Presentation Layer (Client)                           │
│               React 19 SPA · Vite · Bootstrap 5 · Risk Gauges (Chart.js)         │
└────────────────────────────────────────┬─────────────────────────────────────────┘
                                         │ HTTPS / REST (JSON)
┌────────────────────────────────────────▼─────────────────────────────────────────┐
│                       Enterprise API Gateway & Business Layer                    │
│             Spring Boot 3.4.1 (Java 21) · Spring Security · JWT · JPA            │
└──────────────────┬───────────────────────────────────────────────┬───────────────┘
                   │ JDBC (HikariCP)                               │ RestClient (HTTP)
┌──────────────────▼───────────────────┐       ┌───────────────────▼───────────────┐
│           Persistence Tier           │       │     Inference Microservice        │
│        GCP Cloud SQL (MySQL 8.0)     │       │   FastAPI · Python 3.12 · XGBoost │
│   (Users, Predictions, Audit Logs)   │       │      (GCP Cloud Run Container)    │
└──────────────────────────────────────┘       └───────────────────────────────────┘

Project Structure

MirAI/
├── Dataset/                              # ADNI baseline cohort data (CSV files)
├── Docs/                                 # Clinical documentation & interactive assets
│   ├── Deployment_steps.md               # Step-by-step GCP Cloud Run deployment guide
│   ├── MirAI_Methodology_Evolution.md    # In-depth clinical ML audit & methodology evolution
│   ├── MirAI_modelling.ipynb             # Training, cross-validation, calibration & SHAP notebook
│   ├── index.html                        # Interactive methodology & flowchart presentation portal
│   └── output.png                        # Pipeline diagrams and visual assets
├── Major Project/                        # Academic research publication drafts
│   ├── ADpaper.tex                       # Complete LaTeX research paper
│   └── Ai check audit.pdf                # Originality & AI audit reports
├── WebApp/                               # Full-Stack Application Ecosystem
│   ├── backend/                          # Java 21 / Spring Boot 3.4.1 backend service
│   │   ├── pom.xml                       # Maven build configuration
│   │   └── src/main/java/com/mirai/      # REST Controllers, Entities, Services & Security
│   ├── frontend/                         # React 19 / Vite SPA frontend
│   │   ├── package.json                  # Frontend dependencies
│   │   └── src/                          # Assessment wizard, Dashboard, History, Hooks
│   ├── Dockerfile                        # Multi-stage container definition
│   └── PRD.md                            # Product Requirements Document
├── model deployment/                     # Python ML Inference Microservice
│   ├── Dockerfile                        # Cloud Run container configuration
│   ├── app.py                            # FastAPI REST endpoints
│   ├── requirements_api.txt              # Microservice dependencies
│   └── test_api.py                       # Automated API integration tests
├── models/                               # Serialized model artifacts (.joblib) & schemas
│   ├── mirai_features.json               # Input feature schemas
│   ├── mirai_stage1_model.joblib         # Stage 1 Demographic classifier
│   ├── mirai_stage2_model.joblib         # Stage 2 Genetic classifier
│   └── mirai_stage3_model.joblib         # Stage 3 Calibrated Liquid Biopsy classifier
└── README.md                             # Project documentation

Getting Started

Follow these steps to set up and run MirAI locally.

Prerequisites

  • Java Development Kit (JDK): Version 21+
  • Node.js & npm: Node 20+ and npm 10+
  • Python: Version 3.12+
  • MySQL: Version 8.0+ (or active GCP Cloud SQL instance)

Local Installation & Setup

1. Clone the Repository

git clone https://github.com/Varghese778/MirAI-A-Machine-Learning-Cascade-for-Liquid-Biopsy-Based-Alzheimer-s-Disease-Triage.git
cd MirAI-A-Machine-Learning-Cascade-for-Liquid-Biopsy-Based-Alzheimer-s-Disease-Triage

2. Start the Python ML Inference Microservice

cd "model deployment"
python -m venv venv
# Activate: venv\Scripts\activate (Windows) or source venv/bin/activate (macOS/Linux)
pip install -r requirements_api.txt
uvicorn app:app --reload --port 8000

The API documentation will be available at http://localhost:8000/docs.

3. Start the Spring Boot Backend

cd ../WebApp/backend
./mvnw spring-boot:run

The backend server will start on http://localhost:8080.

4. Start the React Frontend Application

cd ../frontend
npm install
npm run dev

Open your browser and navigate to http://localhost:5173.


API Reference

Predict Patient Risk: POST /predict

Request Body:

{
  "AGE": 74.5,
  "PTGENDER": "Female",
  "PTEDUCAT": 16.0,
  "APOE4": 1.0,
  "AB42_F": 450.2,
  "AB40_F": 7200.0,
  "AB42_AB40_F": 0.0625,
  "pT217_AB42_F": 0.085,
  "NfL_Q": 32.4,
  "GFAP_Q": 185.0
}

Response (200 OK):

{
  "prediction": "High Risk",
  "confidence": 0.8124,
  "stage1": {
    "risk_category": "Moderate Risk",
    "risk_probability": 0.542
  },
  "stage2": {
    "risk_category": "Moderate Risk",
    "risk_probability": 0.618
  },
  "stage3": {
    "risk_category": "High Risk",
    "risk_probability": 0.8124
  },
  "timestamp": "2026-08-21T14:00:00.000Z"
}

Evaluation & Results Summary

Stage Input Data Classifier Validation AUC (95% CI) Significance vs Previous Tier
Stage 1 Demographics (AGE, PTGENDER, PTEDUCAT) Logistic Regression 0.6899 (0.6407–0.7376) Baseline
Stage 2 Stage 1 + APOE4 Genotype Logistic Regression 0.7030 (0.6550–0.7510) $p = 0.014$
Stage 3 Stage 2 + UPENN Plasma Liquid Biopsy Panel Calibrated XGBoost 0.8079 (0.7680–0.8478) $p < 0.001$ ($\Delta\text{AUC } +0.105$)

Roadmap

  • [x] Multi-stage ML cascade with leakage-controlled validation
  • [x] Isotonic probability calibration & Decision Curve Analysis (DCA)
  • [x] Microservice architecture deployment on Google Cloud Run
  • [x] Full-stack clinical assessment portal with JWT authentication
  • [ ] Cohort-level batch triage view for hospital neurologist clinics
  • [ ] Automated volumetric MRI feature ingestion (ADNI / OASIS cohorts)
  • [ ] HL7 / FHIR standard EHR integration adapter

License & Disclaimer

Distributed under the MIT License. See LICENSE for more information.

Medical Disclaimer: MirAI is an academic research prototype and Clinical Decision Support System (CDSS). It is not a certified medical device and is not intended for primary clinical diagnosis or treatment planning. All risk stratifications require validation by a certified neurologist.


Acknowledgements

  • Alzheimer's Disease Neuroimaging Initiative (ADNI): Data collection and sharing was funded by the ADNI (National Institutes of Health Grant U01 AG024904) and DOD ADNI (Department of Defense award number W81XWH-12-2-0012). Full protocols and investigator listings are available at adni.loni.usc.edu.
  • GE HealthCare Precision Care Challenge 2026: Developed as a solution for the AI-Driven Prioritization System for Early Alzheimer’s Diagnostic Pathways track.

Built by Sharon Varghese · Vel Tech R&D Institute of Science and Technology

Built With

  • adni
  • biomarkers
  • cloudrun
  • fastapi
  • gcp
  • react
  • shap
  • springboot
  • xgboost
Share this project:

Updates