Inspiration

When software engineers write broken code, compilers throw stack traces that isolate the exact line where execution failed. When a student fails a physics or engineering problem, traditional education hands them a red cross, an arbitrary score like 60%, and tells them to "study harder."

Standard study platforms make this worse by rewarding surface-level pattern recognition. Decades of cognitive science reveal how misleading this is:

  • The False-Positive Trap: David Treagust’s 1986 research demonstrated a 52% false-positive rate in conventional multiple-choice tests—more than half of students who pick the correct answer do so based on flawed causal logic.
  • The Explanatory Illusion: Rozenblit and Keil (2002) documented an 80%+ collapse in explanation when learners are asked to explain physical mechanisms, masking their confusion behind textbook vocabulary (Richard Feynman’s Nominal Fallacy).

By the time a student struggles with an advanced topic like Capacitive Reactance, the issue is rarely that specific chapter. The real failure almost always traces back to a hairline fracture in an earlier prerequisite—like Conservation of Charge or Phase Relationships—that traditional quizzes failed to catch. I built Tessera to act as a conceptual debugger: turning static textbooks into causal prerequisite graphs, intercepting lucky guesses, and tracing downstream failures back to the broken foundation.


What it does

Tessera transforms static textbook chapters into an interactive, diagnostic learning environment:

  1. Student Onboarding & Analogical Anchoring: The landing page introduces the workflow through a guided sticky-scroll overview. Upon entering the studio, students select an intuitive personal domain (like Music Production, Gaming, or Cooking) to serve as a cognitive bridge.
  2. Automated Prerequisite Graphs: Uploading a textbook PDF compiles the text into an interactive Directed Acyclic Graph (DAG). Concepts appear as carved stone slabs with 3-pip pebble mastery gauges, with downstream topics locked to prevent cognitive overload.
  3. In-Depth Exploration via Tessella: Clicking any node slides open the node details panel and launches Tessella, the built-in assistant. Tessella explains concepts through the student's chosen hobby analogy (using Dedre Gentner's Structure-Mapping Theory) and answers technical questions using verified citations retrieved directly from the textbook via vector RAG.
  4. Two-Tier Diagnostics & Anti-Jargon Guard:
    • Tier 1 (Prediction MCQ): Tests the physical outcome (What happens?).
    • Tier 2 (Causal Driver MCQ): Tests the underlying mechanism (Why does it happen?). The distractors represent documented scientific misconceptions mapped to upstream prerequisites.
    • Anti-Jargon Challenge: If a student passes Tier 2, they face a 45-second sprint requiring them to describe the mechanism using plain, physical action verbs rather than memorized buzzwords.
  5. Bayesian Mastery Tracing: Latent understanding is tracked deterministically using Bayesian Knowledge Tracing (BKT), separating real mastery from lucky guesses and careless slips.
  6. Fault-Tree Root Cause Analysis (RCA): If a student selects a misconception distractor in Tier 2, Tessera initiates reverse DAG traversal to locate the upstream prerequisite linked to that failure. It serves a 30-second timed micro-probe on that earlier node:
    • If failed: The upstream prerequisite's mastery drops, its state changes to Fragile (marked with a cracked visual badge), and the student is routed back to repair that foundation before moving forward.
    • If passed: The failure is classified as an isolated slip ($P(S)$), keeping the upstream node intact and offering a targeted hint on the current node.

How we built it

Tessera is built as a decoupled full-stack application that pairs deterministic mathematical scoring with low-latency LLM inference:

  • Frontend: Built with React.js, Node.js, and Tailwind CSS. It manages the dynamic DAG canvas, animated clue trails during backtracking, and responsive inspection drawers.
  • Backend: Powered by FastAPI (Python), handling asynchronous requests, PDF text extraction, topological cycle detection, and Bayesian computations.
  • Database & Vector Search: Supabase (PostgreSQL) stores graph lineages, student progress, and Bayesian mastery parameters. Upstash Vector provides serverless embeddings for grounded, hallucination-free textbook RAG citations.
  • Dual-Model Inference:
    • GPT OSS 120B: Parses raw textbook chapters, synthesizes Gentner-style hobby analogies, and generates atomic concept nodes and dependency edges.
    • Groq (Llama 32B): Handles high-speed natural language evaluation during the 45-second Anti-Jargon sprints, returning semantic validation in under 400ms.
  • Deterministic Bayesian Knowledge Tracing (BKT): Handled purely in Python arithmetic rather than by an LLM. Given prior mastery $P(L_t)$, guess probability $P(G)$, slip probability $P(S)$, and transition probability $P(T)$, the posterior update on a correct response is calculated as:

$$P(L_t \mid \text{Correct}) = \frac{P(L_t) \cdot (1 - P(S))}{P(L_t) \cdot (1 - P(S)) + (1 - P(L_t)) \cdot P(G)}$$

The transition step updates latent mastery for the next attempt:

$$P(L_{t+1}) = P(L_t \mid \text{Obs}) + \Big(1 - P(L_t \mid \text{Obs})\Big) \cdot P(T)$$


Challenges we ran into

  1. Extracting Pure DAGs from Cyclical Textbooks: Textbooks frequently cross-reference future material or define ideas cyclically. Early extraction runs often created circular loops that broke graph traversal. I resolved this by enforcing strict Pydantic output schemas and adding topological cycle-detection checks in the backend to ensure clean, acyclic dependency trees.
  2. The Over-Penalizing Backtrack Bug: In early prototypes, RCA triggered whenever a student failed either the Tier 2 question or the Anti-Jargon check. This penalized students who understood the physics but stumbled on vocabulary during a timed sprint. I decoupled linguistic friction from structural misconceptions: Anti-Jargon stumbles are handled locally with instant re-prompting, while reverse DAG traversal is reserved strictly for Tier 2 distractor selections.
  3. Latency in High-Stakes Timed Checks: Evaluating student free-text responses using standard hosted APIs took 3 to 4 seconds, stalling the 45-second Anti-Jargon timer. Switching lexical evaluation to Groq with Llama 32B dropped inference latency to under 400ms, making the interaction feel instant.

Accomplishments that we're proud of

  • Sub-400ms Real-Time Linguistic Verification: Evaluating natural language explanations against strict semantic constraints without lagging the UI.
  • Deterministic Scoring Over Hallucination: Keeping grading entirely out of the LLM's hands by calculating Bayesian probabilities directly in arithmetic.
  • Targeted Fault-Tree Diagnostics: Replacing arbitrary retries with automated graph backtracking that identifies broken upstream prerequisites and assigns actionable Fragile states.
  • Grounded RAG Experience: Integrating Upstash Vector to ensure Tessella’s citations match the uploaded syllabus without hallucinating definitions.

What we learned

  • Cognitive Science Trumps Superficial Gamification: Real engagement comes from intuitive clarity—grounded in Gentner’s analogical mapping, Treagust’s two-tier diagnostics, and Feynman’s nominal fallacy—rather than superficial streaks and point systems.
  • Student Error is a Graph Traversal Problem: Academic struggle is rarely an inability to learn; it is almost always an unaddressed prerequisite gap. Viewing learning as a dependency graph turns grading into a debugging process.
  • Separating Language from Logic is Crucial: Distinguishing between vocabulary recall and causal understanding is essential when evaluating technical comprehension with AI.

What's next for Tessera

  1. Multi-Modal Document Ingestion: Expanding the pipeline to ingest lecture audio recordings and presentation slide decks alongside textbook PDFs to build unified course graphs.
  2. Collaborative Prerequisite Heatmaps: Providing aggregated knowledge graphs for educators to pinpoint shared prerequisite bottlenecks across entire cohorts.
  3. Spaced-Decay Diagnostic Probes: Incorporating Ebbinghaus forgetting curves into the Bayesian knowledge model, triggering quick micro-probes on past nodes to prevent prerequisite decay over the semester.

Built With

Share this project:

Updates

Submission history