Inspiration
There is a famous set of studies in which clinicians were given a screening scenario: a rare condition, a decent test, a positive result. Most estimated the chance of disease at 70 to 90 percent. The correct answer was under 10. That gap between what a positive result feels like and what it actually means is one of the best-documented failures in health communication, it affects patients and professionals alike, and almost every risk tool on the web quietly sidesteps it by showing risk before a test or test accuracy on its own, never the two combined. Combining them properly, for a person rather than a population, is the whole idea here.
What it does
You enter ten things you already know about yourself, with no laboratory value among them: age band, BMI, sex, general health, diagnosed high blood pressure or cholesterol, heart disease history, activity, mobility, smoking history. A model turns that into your pre-test probability of type 2 diabetes.
Then you pick a screening test (HbA1c, fasting glucose, or a random glucose screen), keep or adjust its sensitivity and specificity, and choose the result you are holding: positive or negative.
The answer comes back as one hundred people. For a 50-year-old woman with BMI 27 and nothing diagnosed, a positive HbA1c means about 36 of 100 people like her actually have diabetes; 64 are false alarms. For a 70-year-old with BMI 34, high blood pressure and high cholesterol, the identical positive means 92 of 100. Same paper, different person, opposite meaning.
Flip to negative and you see the number nobody is ever shown: with a test that misses half of true cases, half the people it clears still have the condition.
How we built it
The pre-test model is a logistic regression fitted on the CDC's 2015 Behavioral Risk Factor Surveillance System sample, 253,680 responses (UCI Machine Learning Repository, id 891), using scikit-learn. Coefficients are exported into the page, so the app itself is three static files: no build step, no backend, no network call, nothing to install.
The result is rendered as a 100-person icon array drawn in pure CSS. Natural frequencies are read more accurately than percentages, which is why the answer is drawn before it is stated.
Calibration mattered more than accuracy here, and the app says so out loud. If the pre-test probability is off by a factor of two, the final answer is off by roughly the same factor. So the model page reports calibration first: on 76,104 held-out responses the model never saw, mean calibration error is 0.013 across ten risk bins, with the full predicted-versus-observed table rendered in the app. ROC AUC is 0.819, and it is deliberately listed second.
Challenges we ran into
The hardest decision was what not to claim. Published sensitivity and specificity for the same test disagree between studies and populations, and a tool that silently picks one number is lying by omission. So both stay adjustable, and the app prints how far the answer moves across a plausible range: the 36% example above swings between 21% and 73%. Showing that band honestly, without making the tool feel useless, took the most design iterations.
Accomplishments that we're proud of
The model is genuinely well calibrated on held-out data, which is the one property this use case cannot live without. And the whole thing runs anywhere a browser runs: the demo is a static page, reproducible from a single training script that fetches the public dataset itself.
What we learned
Absence of evidence reads like evidence to people. A missing confirmation, a cleared test, a "negative" — each feels definitive and often is not. Turning that into one hundred small figures does more than any paragraph of explanation we tried.
What's next for One Hundred People Like You
More conditions and tests, since the machinery generalises to any screening context with published test characteristics. A printable one-page summary a person could bring to an appointment. And validation against a non-US cohort, which the limits section currently flags as missing.
Log in or sign up for Devpost to join the conversation.