Inspiration Most assessment checks whether a student can produce the right answer. It rarely checks whether they understand it well enough to defend it. We kept coming back to the protégé effect — the finding that the fastest way to learn something is to teach it — and asked what a quiz would look like if it took that seriously instead of bolting it on as a bonus activity. The answer wasn't a quiz at all: it was a debate, where an AI persona commits, in good faith, to a specific and plausible wrong belief, and the only way to "pass" is to argue it out of that belief for real.
What it does Protégé Gauntlet generates a specific, plausible misconception in a domain the student picks (physics, software architecture, finance, law, math), then puts The Challenger — an AI debater — in the wrong, confidently defending it. The student argues back; a separate Judge Agent grades, after every message, whether the argument actually addresses the flawed mechanism (not just the textbook answer), and a live Belief Meter visibly cracks as real objections land. A Deterministic Guardrail checks every debater line before it's shown, so the persona can't break character or concede early. A Cognitive Dependency Graph renders the belief structure live and turns green node by node as it's dismantled. For technical concepts, a built-in code sandbox offers a second path to victory: run an O(n²) sort at real scale and watch it time out, and that measured evidence alone can force a concession — no essay required. Every win is logged to a persistent ledger that drives a Reasoning Depth Score, a before/after "misconception erasure" card, adaptive difficulty (misconceptions get harder and subtler the more you defeat), rank titles, achievement badges, and a cohort leaderboard.
How we built it FastAPI + SQLite backend, React (Vite + Tailwind) frontend. The backend is split into strictly independent modules — misconception generator, debater agent, judge agent, guardrail, graph builder, sandbox executor, ledger — coordinated by a single orchestrator state machine, so a chat argument and a sandbox code execution both resolve through the exact same concession logic. Every agent has a fully deterministic mock implementation as well as a real-LLM one (any OpenAI-compatible endpoint), so the whole thing runs standalone with zero API keys and degrades gracefully back to mock logic if a live call ever fails. The live graph streams over SSE; the code sandbox proxies to Piston by default with a local subprocess fallback for offline dev. Deployed as a Render Blueprint, both services wired together entirely through environment variables.
Challenges we ran into Making a mock judge — plain keyword overlap, no model call — actually fair was the hard part. Raw overlap can be fooled by a student who just restates the flawed belief in its own vocabulary, so we built a stance check that tells "refuting it" apart from "reasserting it," plus same-root and WordNet-backed synonym matching so a correct answer phrased differently isn't penalized. We also found a real bug live: the matcher for the flagship code-sandbox concept used the key "sort", which never matches the word "sorting" — silently routing the exact demo scenario from our own docs to a worse fallback. And the offline sandbox's resource-limiting code used POSIX-only APIs that crashed outright on Windows, which needed graceful guarding rather than a rewrite.
Accomplishments that we're proud of Two completely different kinds of evidence — written argument and executed code — resolve through one identical state machine, with no special-casing anywhere in the belief-meter or concession logic. The empirical-evidence path is genuinely satisfying: submitting code that just times out at scale is, on its own, enough to win. And the entire multi-agent pipeline — generator, debater, judge, guardrail — runs convincingly with zero API keys, including real stance-detection and genuine synonym tolerance, not just a keyword-matching stand-in.
What we learned A deterministic heuristic can substitute for real understanding much further than expected, as long as you're deliberate about what signal you're actually measuring — stance, not just overlap — and honest about where its ceiling is. Keeping the debater and the judge as two genuinely separate agents (one persuadable-but-stubborn, one dispassionate) is what stops the persona from ever grading its own homework. And a one-word mismatch in a matching key can silently break a flagship feature, so it's worth testing the exact path a real user takes, not just the logic underneath it.
What's next for Protégé Gauntlet A live PvP "cross-examination" mode — two students racing to defeat the same misconception — is the highest-impact addition on the table, deliberately deferred for scope. Beyond that: real authentication instead of a slugified display name, moving live session state to Redis for true multi-instance deployment, expanding the Fallacy Matrix into the domains already staged as "coming soon" (Medicine & Anatomy, Corporate Law), and eventually defaulting to the real LLM judge with the deterministic path kept as the offline fallback it already is.
Log in or sign up for Devpost to join the conversation.