-
-
Human-in-the-Loop Review — Record reviewer decisions, escalation actions, and reasons for flagged predictions.
-
Batch Drift Monitoring — Detect distribution changes using MMD statistics and distinguish shift alerts from individual safety decisions.
-
Safety Workbench — Compare PASS and BLOCK decisions to see why high model confidence does not guarantee safety.
-
Decision Trail — Inspect recorded analysis events, reviewer actions, supporting evidence, and hash-chain verification.
-
What-if Scenario Laboratory — Compare model safety decisions and simulated reviewer workload under different operating conditions.
-
Automatic Drift Hold — See how detected distribution changes can pause processing and require human acknowledgement.
-
Live 3D Workflow Twin — Explore connected AI processing stages from image intake to guardian decisions and retained outputs.
-
Interactive 3D Digital Twin — Visualize the medical AI workflow, processing stations, and live operational controls.
-
Evidence Room — Compare model accuracy, confidence, and error rates across seven controlled image conditions.
MedShift Guardian 3D
See the flow. Detect the risk. Keep humans in control.
INSPIRATION
Medical AI can perform well during testing but become unreliable when image quality or data distributions change. One serious problem is that an AI model can remain highly confident even when its predictions are wrong.
We built MedShift Guardian 3D to explore how medical AI risks can be detected, explained, and presented to human reviewers before automated decisions are accepted.
WHAT IT DOES
MedShift Guardian 3D is an interactive medical AI safety research platform that combines AI inference, monitoring, 3D visualization, evidence analysis, and human oversight.
The platform includes seven major capabilities:
Interactive 3D Digital Twin: Visualizes image intake, AI inference, safety monitoring, human review, and output decisions in an interactive 3D environment.
Safety Workbench: Tests medical benchmark images under different image-quality conditions and generates PASS, REVIEW, or BLOCK monitoring decisions.
Batch Monitor: Uses Maximum Mean Discrepancy (MMD) and permutation testing to identify distribution changes across batches of medical images.
Evidence Room: Presents model accuracy, confidence-related errors, false alerts, prediction coverage, and monitoring performance.
Human-in-the-Loop Review: Allows human reviewers to inspect flagged jobs and record decisions with supporting reasons.
Decision Trail: Stores analysis and review events using SQLite and hash-linked records for local integrity verification.
Resource Context: Supports clearly labeled operational and planning records while distinguishing unknown information from verified observations.
HOW WE BUILT IT
We developed the backend using Python and built the interactive frontend with HTML, CSS, JavaScript, and Three.js.
The research model uses PneumoniaMNIST, a public pediatric chest X-ray benchmark from MedMNIST. Its classification pipeline combines Principal Component Analysis (PCA) with a small neural network.
The monitoring system evaluates image quality, reference novelty, prediction ambiguity, and batch-level distribution differences.
The 3D Digital Twin visualizes the software workflow and its recorded states. It does not represent a connected physical hospital.
SQLite stores analysis records, reviewer actions, and supporting audit information. The application is hosted on Render, and its source code is available on GitHub.
CHALLENGES WE FACED
Our first challenge was demonstrating that high model confidence does not necessarily mean a correct prediction.
The second challenge was distinguishing distribution shift from actual model failure. A statistical distribution alert alone cannot establish clinical harm.
The third challenge was balancing risk detection against false alerts and unnecessary human-review workload.
We also worked to make technical monitoring results understandable through a visually interactive interface rather than relying entirely on conventional dashboards.
ACCOMPLISHMENTS
Our recorded benchmark evaluation produced the following results:
Clean-image classification accuracy: 85.4%.
Accuracy under controlled overexposure: 63.8%.
High-confidence errors under overexposure: 213.
Clean false-alert rate for the combined monitoring policy: 18.8%.
These results demonstrate why AI monitoring systems must report both detected risks and remaining limitations.
We also integrated the 3D workflow, safety evaluation, batch monitoring, human review, and evidence reporting into one runnable web application.
WHAT WE LEARNED
We learned that trustworthy medical AI requires more than prediction confidence.
Monitoring systems must distinguish unusual input data from confirmed model failure, communicate uncertainty clearly, and preserve evidence for human review.
We also learned that monitoring accuracy alone is insufficient. False alerts, missed errors, and prediction coverage must be evaluated together.
WHAT'S NEXT
Our future goals include evaluating the system against independent scanner and hospital distribution shifts, improving monitoring accuracy, reducing unnecessary alerts, and conducting usability studies with medical AI reviewers.
We also plan to investigate stronger validation methods and more robust audit infrastructure.
MedShift Guardian 3D is an educational research prototype. It is not clinically validated, is not connected to a live hospital, and must not be used for patient-care decisions.
PROJECT LINKS
Live Website: https://medshift-guardian-3d.onrender.com
GitHub Repository: https://github.com/sadia-zannat/medshift-guardian-3d
Dataset: PneumoniaMNIST / MedMNIST v2
Built With
- css3
- github
- html5
- javascript
- numpy
- pneumoniamnist
- python
- render
- scikit-learn
- sqlite
- three.js

Log in or sign up for Devpost to join the conversation.