Why it exists
Fixed flashcard schedules can test a familiar item too soon and a difficult one too late. The same interval is a poor fit for both the word "a" and an irregular Spanish subjunctive, and it ignores how one learner's history differs from another's. Half-Life turns that mismatch into a prediction problem.
What Half-Life does
Duolingo released 12.9 million language-learning review logs. Half-Life trains on 10,243,562 of them. It makes an item's half-life a function of that learner's recall history and an item-specific weight. If the half-life is h days, predicted recall after d days is 2^(-d/h). Shared model weights are fitted with stochastic gradient descent over the review stream.
In the browser demo, learners compare predicted recall curves with a fixed Leitner schedule. They can also add a vocabulary card with their own answer, reveal it, and mark it recalled or forgotten. Review counts and history stay in the same browser across restarts. Displayed review times are model estimates; cards can be reviewed sooner.
Held-out evidence
Learners were split 80/10/10 by user ID. Hyperparameters were chosen only on the validation learners, and the test learners were scored once. No learner appears in more than one split.
Across 1,279,602 held-out reviews, Half-Life reaches 0.1208 mean absolute recall error, compared with 0.1727 for an average-recall baseline, 0.2224 for Leitner, and 0.4367 for Pimsleur. That is 46% less error than Leitner. Half-life error is 116.3 days, versus 147.4 for Leitner and 154.3 for Pimsleur.
Leitner has higher AUC, 0.5496 versus 0.5405. Lower mean absolute error is not a separate calibration test or evidence of improved learning. Duolingo chose the review intervals in these observational logs, which can affect the fitted coefficients. The demo now labels intervals as predictions and displays this limitation.
How it is built
Training uses Python's standard library, no GPU, no API key, and about six minutes on a laptop CPU. The model exports small JSON weights. The public demo loads them and runs every prediction in the browser without a server call after load. Failed data loads show an error and keep prediction controls disabled. Study history uses IndexedDB transactions; duplicate review attempts count once. Answers and review events are not uploaded or used to refit the model.
Repository: https://github.com/HyunsikParker/half-life Demo: https://hyunsikparker.github.io/half-life/
Team, AI tools, and data
This is a solo entry. Claude by Anthropic assisted with much of the Python and JavaScript and drafted the initial description. Codex assisted with the study flow, browser storage, error handling, and reporting corrections. The linked repository produces the held-out results above. The implementation follows Settles and Meeder's 2016 ACL paper, "A Trainable Spaced Repetition Model for Language Learning."
The dataset is Duolingo Learning Traces, CC BY 4.0, doi:10.7910/DVN/N8XJME. The build downloads it; the repository does not redistribute it.
Log in or sign up for Devpost to join the conversation.