Inspiration

The Feynman technique learning by teaching is one of the most effective study methods. But teaching a passive listener doesn't test understanding. Teaching someone who actively resists with a specific wrong belief forces you to explain why the misconception is wrong, not just recite the correct answer. That's the gap Reverse Socratic fills.

What it does

Reverse Socratic flips AI tutoring upside down. Instead of the AI teaching you, you teach the AI but the AI plays a confused student who holds real, documented misconceptions from science education research. Your job is to explain the concept well enough to actually correct those wrong beliefs. A second AI then evaluates your teaching and scores you 0–100.

How a session works:

  1. Pick a concept electricity, photosynthesis, natural selection, why the sky is blue, or seasons. Each has 2–3 real misconceptions drawn from science education literature.
  2. The AI student opens with its wrong belief and pushes back 1–2 times with genuine confusion. It only accepts a correction when your explanation addresses the root of the misconception.
  3. A separate evaluator AI reads the full conversation and determines which misconceptions were actually corrected (not just mentioned). You get a score, per-misconception verdicts, and specific feedback.

How we built it

Two distinct Gemini agents run in concert, this is not a single chat wrapper:

  • The confused student (Gemini 3.6 Flash, streaming via SSE) initialized with a system prompt encoding the concept's misconceptions and a "resist until taught properly" behavior. It pushes back against shallow explanations and only accepts corrections that dismantle the root of the wrong belief.
  • The evaluator (Gemini 3.6 Flash, JSON structured output) reads the full conversation and uses a strict response schema to return per-misconception verdicts, an overall score, strengths, and improvements. It distinguishes "mentioned the misconception" from "actually corrected it."

Stack: Next.js 16 (App Router, TypeScript), Tailwind CSS v4 with a Duolingo-inspired design system, GSAP for hero animations and animated SVG concept icons, Thinking Orbs for AI status indicators (working, listening, solving), deployed on Vercel. No database, no auth, fully stateless.

Challenges we ran into

  • Getting the AI to resist convincingly too passive and it accepts any answer (defeats the purpose); too stubborn and it never accepts correct explanations (frustrating). Tuning the system prompt and temperature (0.6) to hit the sweet spot took several iterations.
  • Streaming with graceful truncation Gemini's maxOutputTokens cap can cut responses mid-sentence. We check finishReason for MAX_TOKENS and append an ellipsis so the UI degrades gracefully instead of showing a broken cutoff.
  • Evaluator accuracy the evaluator had to distinguish "teacher mentioned the misconception" from "teacher actually corrected it." Structured output with a strict response schema and a prompt that defines "cleared" as requiring root-cause explanation solved this.

Accomplishments that we're proud of

  • The two-agent architecture works: the student agent stays in character and resists convincingly, and the evaluator agent accurately distinguishes real corrections from surface-level mentions.
  • The misconception badges flip in real time as the conversation progresses, giving immediate visual feedback that the teaching is working.
  • Five concepts with 15 researched misconceptions, all drawn from real science education literature, not invented.

What we learned

  • Prompt engineering for persona consistency is harder than it looks. The confused-student persona needed explicit rules for when to accept vs. resist, how many times to push back, and how to show genuine understanding when corrected.
  • Structured output (JSON schema) is far more reliable than prompt-asking for JSON when you need a strict evaluation format.
  • Streaming UX matters. Adding "Thinking..." and "I am listening attentively..." indicators with animated orbs made the AI feel alive rather than like a loading spinner.

What's next for Reverse Socratic

  • More concepts across physics, biology, chemistry, and math
  • Difficulty levels: a stubborn student who pushes back more vs. one who accepts corrections faster
  • Spaced repetition: track which misconceptions you've successfully taught and revisit ones you struggled with
  • Classroom mode: teachers assign concepts and see aggregate scores across students

Built With

Share this project:

Updates

Submission history