Inspiration Benjamin Bloom's 1984 "2-Sigma Problem" showed that one-on-one tutoring outperforms classroom instruction by two standard deviations — yet has always been economically impossible to scale. Every tutoring app we looked at was really an answer-generating chat wrapper: it told you the answer and called it learning. That defeats the entire point — the value of a tutor isn't the answer, it's the misconception diagnosis that happens before the answer. We set out to build the thing that closes 2-sigma at zero marginal cost: a multi-agent tutor that never reveals the answer, runs on free tiers and in-browser compute, and keeps student data local by default. What it does Aura 9.0 is an adaptive AI tutor built on three personas that can't break character:

  • Socratic Tutor — diagnoses your misconception (from a closed five-category taxonomy: causal-confusion, overgeneralization, definitional-gap, procedural-slip, prerequisite-missing) and asks exactly one guiding question. It never states the answer, even when you beg it to — jailbreak resistance is a tested contract, not a hope.
  • Curious Novice — runs Reverse Learning: it pretends to know nothing and interrogates you, probing vague explanations for concrete examples and jargon for plain-language definitions, then grades your transcript into a mastery report.
  • Knowledge Graph + BKT engine — every concept is a node with prerequisite edges; Bayesian Knowledge Tracing updates mastery from every answer, coloring the graph green/red and routing your next tutoring session to exactly the concept you're weakest on. It ingests your PDFs (Supabase Storage → chunking → bge-small embeddings → pgvector), grounds every answer in retrieved content (if retrieval comes back empty, it says so), runs Whisper STT and a 1B-parameter LLM entirely in your browser via WebGPU when available, and falls back silently to server models when it isn't. All three personas, the WebGPU path, the Groq→Gemini rate-limit path, and a collaboration space with WebRTC peer audio and a mediation agent — one product, zero subscription. How we built it A monorepo: Next.js 14 App Router frontend, FastAPI backend, shared types package. Supabase for auth/Postgres/pgvector, Neo4j Aura for the knowledge graph, Groq (Llama-3-70B) and Gemini 1.5 Flash for the three LLM agents, WebLLM + Transformers.js for in-browser inference, tldraw for the whiteboard, D3 for the graph. The build itself was agentic — the repo runs under a "constitution" (AGENTS.md) and a phase-by-phase plan (plan.md) that a coding agent executes through a tiered MCP toolchain: Context7 fetches version-pinned SDK docs before every external API call, Vitest runs after every code write, Playwright drives every UI flow, Chrome DevTools verifies console/network/Core Web Vitals after every change, and a phase gate (full perf trace + Lighthouse + screenshot diffing) closes every phase. Prompt templates carry JSON-schema contracts, and an answer-leakage regression suite runs at every gate from Phase 2 until launch, forever. Challenges we ran into
  • Answer leakage as a bug class. The product's entire premise is that the tutor never reveals the answer. We had to build an adversarial test suite — fixtures where the student says "just tell me," "skip the questions," "I already know, confirm it" — and we treat a single leak as a P0, not a flaky test.
  • Fallback chains as first-class code paths. WebGPU absent → Groq; Groq 429 → Gemini; Neo4j quota → Postgres mirror. Each had to be tested under forced failure, not just wired theoretically. A rate-limit cascade mid-demo is a silent product killer.
  • LLM personas under pressure. The Reverse Learning interrogator has two documented stress cases — the student who rambles and the student who refuses to answer — and it must never signal it already knows the answer. That's a state-machine problem and a prompt problem.
  • "Confidently wrong" SDK code. Fast-moving APIs (WebLLM, pdf.js, Supabase client, streaming Groq) made hallucinated method names a real risk; we stopped trusting memory and made live docs lookup mandatory before every external API call.
  • The dev tooling itself. Even the mundane things bit back — a dev server silently refusing to serve a new route group (login was a 404 until we forced a recompile), a parser reading cmd.type when the API says cmd.method, React StrictMode double-firing effects and duplicating logs. Accomplishments that we're proud of
  • The Socratic contract holds. Every prompt-template change must pass the leakage suite before merge — small wording tweaks are exactly how leakage regressions sneak in, so there are no exceptions.
  • Zero-marginal-cost architecture, honored. WebLLM and Whisper run in-browser; the fallback models are Groq/Gemini free tiers; no paid dependency entered the stack without explicit human sign-off.
  • Privacy by architecture, not by policy. The service role key never reaches the client bundle; every user-data table ships with RLS and a negative-access test proving user A can't read user B's rows.
  • A working algorithm visualizer — the Smart Class whiteboard parses real algorithm code (merge sort, Dijkstra) into 136 step-by-step visual states with multi-element highlighting, step controls, and speed control. What we learned
  • Verification is the product. The retry ceiling in our build loop (3 failed patches → re-plan, 5 → stop and escalate to a human) was the single most valuable process decision: an autonomous loop without a ceiling is just an infinite thrash loop.
  • A closed taxonomy beats an open-ended field. Forcing the Evaluator into five misconception categories prevents taxonomy drift and makes every downstream decision (tutor prompt, BKT update, graph coloring) tractable.
  • Determinism you can test beats cleverness you can't. Hand-calculated BKT fixtures, strict Mermaid-syntax validators, mocked navigator.gpu fallback tests — the boring tests caught the real bugs.
  • Graceful degradation is a feature. Students don't care which model answered; they care that it answered. Silent, tested fallbacks make the product feel more reliable, not less. What's next for Aura 9.0
  • Phase 4 & 5 completion — WebRTC collaboration spaces with in-browser Whisper and the 30-second mediation loop, then full UI polish and fallback hardening (forced-429 → Gemini verified end-to-end).
  • Cross-feature knowledge graph sync — a wrong answer in a Socratic session, a low mastery report, and a collaborative-session misconception all write to one consistent BKT state per concept, so the graph is always the single source of truth.
  • The 500-node stress test as a standing regression — the graph must render without jank at scale, not just on demo data.
  • Live deployment with Sentry watch — every phase gate already gates on full verification; shipping with continuous error monitoring is the next natural step.
  • Longer horizon — tune BKT parameters on real student data, add OAuth providers, and scale the collaboration rooms beyond two peers.

Built With

  • bkt
  • cypher
  • llm-agents
  • neo4j
  • postgresql
  • sql
  • supabase
  • typescript-(next.js-14-app-router-+-react)-?-/apps/web-and-/packages/shared-python-(fastapi
Share this project:

Updates

Submission history