Inspiration

Most tutoring tools stop at "here's whether you got it right." That's the symptom, not the disease. I kept thinking about how a real oral exam works: the examiner asks a few sharp follow-up questions until they find the one idea you're actually shaky on, explains that, then checks again. I wanted that experience for gradient descent, where the real misunderstanding is almost always buried one or two prerequisites earlier than where the confusion shows up.

What it does

MindProbe is a diagnostic engine, not a tutor or a quiz generator. Given a fixed prerequisite chain (Slope → Derivative → Gradient → Learning Rate → Gradient Descent), it teaches the whole chain once, then asks the student to explain gradient descent in their own words. From that explanation, it finds the first concept in the chain, from most foundational forward, where understanding actually breaks down, not just whichever concept scored lowest. It asks 3 to 4 Socratic questions targeting that root concept, delivers a short correction naming the misconception directly, then re-tests with a new question designed to be unanswerable while still holding the old misconception. Nothing is revealed until the end, when the student sees the full diagnosis, a before/after concept map, and confirmation of whether the fix stuck, including whether it lifted their understanding of everything downstream.

How I built it

The frontend is Next.js 14 with Tailwind, the backend is FastAPI with SQLAlchemy. The core decision was drawing a hard line between the LLM and the application: the model only ever returns structured assessments, forced through tool calling, and a separate deterministic Python layer is the sole authority on scores, root cause selection, when probing ends, and whether a misconception is resolved. The LLM never touches application state directly. I used Groq's free tier API (openai/gpt-oss-20b) for every generation and assessment call.

Challenges I ran into

Most of the trouble came from the LLM layer being flakier than the deterministic layer around it. gpt-oss-20b is a reasoning model, so it spends part of its token budget on internal reasoning before writing visible output; early on, lesson generation would silently come back empty because the budget ran out first. Structured JSON tool calls hit a similar wall: verbose reasoning fields pushed generations past max_tokens mid object, so Groq rejected the tool call as unparseable JSON. I also hit a schema mismatch where the model returned confidence as a 0 to 1 fraction instead of the 0 to 100 integer the schema demanded. On the frontend, React Strict Mode's double invoked effects caused a race where two lesson requests fired, and whichever resolved second, even an empty one, silently overwrote good content in state. A gauge component built for a full width debug view also broke when reused in a half width panel, causing dials to overlap and overflow off screen. Each of these taught me to treat the model's output as unreliable input, never a guarantee.

Accomplishments that I'm proud of

The LLM and application logic separation held up under real, messy conditions. Every time the model misbehaved, the deterministic scoring and diagnosis logic kept working because it never depended on the model behaving well in the first place. I'm also proud that the delayed reveal flow works end to end against a live API: the student genuinely sees no hint of the diagnosis until the final screen, which is what makes the probing phase feel like an honest assessment instead of a quiz with the answer key showing.

What I learned

Reasoning models need real headroom and lowered reasoning effort for even simple generation tasks, or they burn their budget on invisible thinking and return nothing. Structured output schemas should be defensive about type shape, not just field presence, since models don't always respect an integer constraint even when explicitly told to. A surprising number of "bugs" also turned out to be environment issues, like a stale venv or a version mismatch, that looked identical to code bugs in the traceback.

What's next for MindProbe

Beyond the single gradient descent chain, the natural next step is generalizing the prerequisite chain model to other subjects without hardcoding a new graph each time. I'd also like spaced repetition so a resolved misconception gets checked again rather than assumed permanent, and lightweight accounts so a student's concept map persists across sessions instead of living only in one session ID.

Share this project:

Updates

Submission history