Inspiration Most assessment checks whether a student can produce the right answer. It rarely checks whether they understand it well enough to defend it. We kept coming back to the protégé effect — the finding that the fastest way to learn something is to teach it — and asked what a quiz would look like if it took that seriously instead of bolting it on as a bonus activity. The answer wasn't a quiz at all: it was a debate, where an AI persona commits, in good faith, to a specific and plausible wrong belief, and the only way to "pass" is to argue it out of that belief for real.

What it does Protégé Gauntlet generates a specific, plausible misconception in a domain the student picks (physics, software architecture, finance, law, math), then puts The Challenger — an AI debater — in the wrong, confidently defending it. The student argues back; a separate Judge Agent grades, after every message, whether the argument actually addresses the flawed mechanism (not just the textbook answer), and a live Belief Meter visibly cracks as real objections land. A Deterministic Guardrail checks every debater line before it's shown, so the persona can't break character or concede early. A Cognitive Dependency Graph renders the belief structure live and turns green node by node as it's dismantled. For technical concepts, a built-in code sandbox offers a second path to victory: run an O(n²) sort at real scale and watch it time out, and that measured evidence alone can force a concession — no essay required. Every win is logged to a persistent ledger that drives a Reasoning Depth Score, a before/after "misconception erasure" card, adaptive difficulty (misconceptions get harder and subtler the more you defeat), rank titles, achievement badges, and a cohort leaderboard.

How we built it FastAPI + SQLite backend, React (Vite + Tailwind) frontend. The backend is split into strictly independent modules — misconception generator, debater agent, judge agent, guardrail, graph builder, sandbox executor, ledger — coordinated by a single orchestrator state machine, so a chat argument and a sandbox code execution both resolve through the exact same concession logic. Every agent has a fully deterministic mock implementation as well as a real-LLM one (any OpenAI-compatible endpoint), so the whole thing runs standalone with zero API keys and degrades gracefully back to mock logic if a live call ever fails. The live graph streams over SSE; the code sandbox proxies to Piston by default with a local subprocess fallback for offline dev. Deployed as a Render Blueprint, both services wired together entirely through environment variables.

Challenges we ran into Making a mock judge — plain keyword overlap, no model call — actually fair was the hard part. Raw overlap can be fooled by a student who just restates the flawed belief in its own vocabulary, so we built a stance check that tells "refuting it" apart from "reasserting it," plus same-root and WordNet-backed synonym matching so a correct answer phrased differently isn't penalized. We also found a real bug live: the matcher for the flagship code-sandbox concept used the key "sort", which never matches the word "sorting" — silently routing the exact demo scenario from our own docs to a worse fallback. And the offline sandbox's resource-limiting code used POSIX-only APIs that crashed outright on Windows, which needed graceful guarding rather than a rewrite.

Accomplishments that we're proud of Two completely different kinds of evidence — written argument and executed code — resolve through one identical state machine, with no special-casing anywhere in the belief-meter or concession logic. The empirical-evidence path is genuinely satisfying: submitting code that just times out at scale is, on its own, enough to win. And the entire multi-agent pipeline — generator, debater, judge, guardrail — runs convincingly with zero API keys, including real stance-detection and genuine synonym tolerance, not just a keyword-matching stand-in.

What we learned A deterministic heuristic can substitute for real understanding much further than expected, as long as you're deliberate about what signal you're actually measuring — stance, not just overlap — and honest about where its ceiling is. Keeping the debater and the judge as two genuinely separate agents (one persuadable-but-stubborn, one dispassionate) is what stops the persona from ever grading its own homework. And a one-word mismatch in a matching key can silently break a flagship feature, so it's worth testing the exact path a real user takes, not just the logic underneath it.

What's next for Protégé Gauntlet A live PvP "cross-examination" mode — two students racing to defeat the same misconception — is the highest-impact addition on the table, deliberately deferred for scope. Beyond that: real authentication instead of a slugified display name, moving live session state to Redis for true multi-instance deployment, expanding the Fallacy Matrix into the domains already staged as "coming soon" (Medicine & Anatomy, Corporate Law), and eventually defaulting to the real LLM judge with the deterministic path kept as the offline fallback it already is.

Built With

Share this project:

Updates

Submission history