Inspiration

Most "AI tutor" projects treat a wrong answer as a single bit of information: right or wrong. When you're wrong, they respond by re-explaining the topic — often in the exact same way that didn't land the first time. But that's not how good human tutors work. A great teacher doesn't just say "not quite, here's the answer" — they notice what you were actually thinking when you got it wrong, and address that specific mental model directly. Confusing velocity with acceleration, treating a pointer like a label instead of a numeric address, assuming correlation implies causation — these are specific, nameable, predictable mistakes, not random noise.

We wanted to build a tutor that does that: diagnose the misconception, not just the mistake.

What it does

Faultline takes any topic and breaks it into a small chain of prerequisite concepts, rendered as a vertical cross-section — like geological strata, with foundational ideas as bedrock underneath the concepts that depend on them. Each concept has one diagnostic multiple-choice question. The wrong answers aren't random distractors — each one is engineered, at generation time, to correspond to a specific, named misconception a real student could plausibly hold.

When you pick a wrong answer, Faultline already knows exactly which mental trap you fell into (it was tagged when the question was generated — no guessing required). It shows you that diagnosis immediately, then generates a fresh counter-example on the spot, built specifically to break that mental model rather than restate the correct rule.

Get it right, and that layer of the map turns stable — unlocking the next concept that depends on it.

How we built it

The frontend is Next.js with TypeScript, styled with Tailwind and a light, warm palette with dark, high-contrast small elements (buttons, badges, concept chips). Framer Motion drives two deliberate 3D moments: the concept layers tilt into place like rock strata settling, and the question card does a full 3D flip to reveal the diagnosis panel. A small React Three Fiber scene in the header — a slowly rotating, fractured crystal — reinforces the brand metaphor without turning into a generic 3D flourish.

On the backend, a class-based service layer wraps the Anthropic Claude API behind a Singleton client and an Adapter interface, so the LLM provider is swappable without touching call sites. Every piece of generated content — concept trees and remediations alike — runs through a two-tier cache: Redis for sub-10ms hot reads, Postgres (via Prisma) as the durable source of truth. The key architectural decision: misconception remediations are cached per (question, misconception) pair, not per user or per attempt. The first student who picks a given wrong answer triggers generation; every student after that gets an instant cached response. LLM calls scale with the number of distinct questions and misconceptions in the system, not with usage.

All LLM output is validated against Zod schemas before it's ever persisted, so a malformed generation can't silently corrupt the cache for everyone downstream.

Challenges we ran into

[Fill in honestly once you've built it — e.g. getting the model to generate genuinely distinct, plausible misconceptions rather than one obviously-correct and three obviously-wrong options; tuning the JSON schema prompt until output validated reliably; balancing how much of the 3D/cache architecture was actually feasible solo in the time available.]

Accomplishments that we're proud of

[Fill in — e.g. the deterministic diagnosis layer working end-to-end with zero runtime misclassification; the cache architecture actually eliminating redundant calls measurably; the 3D flip interaction landing well in testing.]

What we learned

[Fill in — e.g. lessons about structured LLM output reliability, prompt engineering for engineered distractors, or working with a new stack component like Upstash/Prisma/R3F for the first time.]

What's next for Faultline

Expanding beyond one question per concept toward a fuller diagnostic battery per topic, adding lightweight accounts so progress persists across devices, and building out a small taxonomy of domain-specific misconception categories (physics, programming, economics) so distractor generation gets sharper with use rather than starting from scratch on every topic.

Built With

Share this project:

Updates