Scores a learner's answer separately from the explanation, detects correct answers supported by incorrect reasoning, and creates a contrastive replay. I used Codex with GPT-5.6 to implement the taxonomy, twenty-case evaluation, tests, and product experience. A future live path can use GPT-5.6 to classify open-ended reasoning and draft replays; the public demo is explicitly a precomputed, tested fixture.

Built With

Share this project:

Updates