Inspiration

I have spent years building tools for teachers, and one complaint never goes away: homework tells you who got it wrong, never why. A stack of wrong answers is a pile of question marks the teacher has to guess through the next morning.

Most educational apps make it worse. Every student gets the same quiz and the same next step, the app never talks back to the teacher, and every session starts from zero. The good tutor at the kitchen table does the opposite: she asks why you picked that, notices what keeps tripping you up, changes how she explains, and remembers you tomorrow.

The Collaborative Partner brief describes exactly that tutor. ClassGame puts her in a 3D classroom on a student's phone, gives her a memory that survives the tab and the mission, and — the part I cared about most — makes what she learns about each student useful to the teacher.

ClassGame is the optional homework that takes notes for the teacher.

What it does

The teacher signs in, creates a class, writes what the lesson covered and what she noticed, and optionally attaches the real worksheets (PDF, images, text). Mia — the Multimodal Intelligence Agent, powered by Gemini — turns that into a mission: 3–4 stages, questions whose wrong options are deliberate traps tied to named misconceptions, step-by-step explanations, a lesson flow map, and a set of knowledge chunks distilled from the teacher's material that Mia will retrieve from while tutoring. The teacher reviews the preview, adjusts the prompt, regenerates (every revision stays on the same mission), and approves. The mission appears in the class in real time.

The student enters the class's six-digit code and a name, picks a character, and walks into a low-poly 3D classroom. A glowing circle appears by the desk the moment the teacher publishes. Mia runs the mission on the chalkboard:

  • She asks every question and decides every next step.
  • When the answer is wrong she does not reveal the solution. She asks why you picked it, with reasons written in the student's voice and generated for that specific wrong answer.
  • She explains in short steps or draws it, depending on what has worked for this student before.
  • The student rates each explanation — helpful, confusing, too easy, too hard — and Mia adapts and says why: "You marked 'confusing' → I'll draw the next one."
  • Close the tab, come back tomorrow: "Welcome back — last time you were on Fractions." Same question, same step, same lines on the board.

Back on the dashboard, the teacher watches the class live from Firestore: who has not started, who is at the desk with Mia right now, who left early, who finished. Then what Mia observed per student: a summary, strengths, areas to revisit, a learning-flow graph of the path they took, and the evidence behind every claim — each answer, each explicit reason, each feedback tap, each tutor turn with Mia's teacher-only observation and the ids of the chunks she retrieved. Across the class: clustered strengths and needs and a tiered plan for the next lesson. One button, Build tomorrow's mission, turns those patterns into the next prompt. That closes the loop.

How we built it

One repository, two Cloud Run services, one database, built alone in three days.

The agent. Go with the Google GenAI SDK on Vertex AI (gemini-3.6-flash), structured JSON output, four generators in one orchestrator: BuildMission, BuildTutorTurn, BuildLearningTrace (after a completed or abandoned attempt) and BuildClassLearningSynthesis. Every generator has a deterministic fallback, so no endpoint fails because of the model, and every output records provider, model, generated_by_ai, fallback_reason and token usage.

One tutor turn. The game dispatches a deterministic tutor locally so the chalkboard never waits; the API turn overwrites Mia's message, the "why" prompt, the explanation steps and the style/difficulty when it lands. Server-side: ground the request against the pinned published question → retrieve the top 5 knowledge chunks for this question and this response → load the student's memory, the attempt history (last 24 observable events, last 12 tutor turns) and the class context → Gemini → a leak guard that discards any output mentioning teacher-only or class context → persist idempotently by event id → return only student-safe fields.

Memory. Three layers in Firestore: the attempt (progress, events, decisions, exact checkpoint, tutor turns, trace), the student (learning_memory: strengths, improvement areas, preferred explanation style, next_adaptation — carried into the next mission) and the class (clustered patterns across students). The game writes a monotonic snapshot 450 ms after every state change, every 30 s while visible, and on pagehide. Cloud Tasks runs finalization on completion and a 120 s abandonment watchdog that still writes a partial trace for a student who gave up; the same queue copies attachments into Cloud Storage. Redis rate-limits generation per teacher, student and IP. Firebase Auth: Google for teachers, anonymous for students.

Frontend. Next.js 16, React 19, TypeScript, Tailwind, Three.js. The classroom is one merged GLB split into connected islands at load time to stretch the room, seat classmates and derive colliders; the six chibi characters are repainted per instance by rewriting vertex colours. Avatars are generated headless in Blender (one shared rig, Draco), which I drove through its MCP server from Claude Code to inspect and iterate. The dashboard uses Firestore realtime listeners, React Flow for the mission and learning graphs, Uploadcare for uploads, PostHog for product and error analytics.

Safety and tenancy. Firestore rules are read-only for browsers — teachers see only their own classes, students only a student-safe projection. Mission generation has a deterministic pre-model gate and a post-model education guardrail; blocked missions cannot be published. Both tutor and trace prompts reason only from observable answers and are forbidden from claiming to read a student's mind. Every durable record already carries organization_id, district_id and school_id.

Challenges we ran into

Keeping the student's screen honest. The first version showed Mia's notebook to the student. It felt agentic and it was wrong: a 12-year-old should not read "confuses numerator and denominator" about themselves. I moved every observation to the teacher side, made the API return a student-safe projection, denied anonymous reads in Firestore rules, and added the leak guard.

Never starting over — without faking it. "Persistent memory" is easy to fake with localStorage. Real persistence meant revision numbers and server-side merge with optimistic concurrency, a checkpoint that includes the lines on the board and the pending decision, revision pinning so a republish cannot swap questions under a running attempt, and a watchdog that finalizes what a student left behind.

Fast and grounded at the same time. A model call per tap makes a chalkboard feel slow. The local deterministic tutor plus the overwriting structured turn keeps the game responsive, and grounding on the server — never the client — closed the door on a browser inventing "expected answers".

Structured output under pressure. Gemini's structured-output endpoint rejected schema keywords I took for granted (maxItems, for one), so a generator that worked in the morning failed after a schema edit in the afternoon. I added a sanitizer that strips unsupported keywords recursively before every call, normalize every blueprint after it, and keep the deterministic fallbacks identical in shape — the UI never sees the difference.

Furnishing a room from one mesh. Seating a classmate in every chair meant recognising chairs and desk tops by their dimensions and measuring the rig's leg pivot. Repainting characters hit 8-bit colour quantisation — near-black hair and near-black shoes became the same colour — solved by matching vertices only against parts that can exist at their height.

Accomplishments that we're proud of

  • Both collaboration loops run live end to end: teacher → mission → student → evidence → next mission, with no manual step in between and no mocked data in the demo.
  • Mia explains every decision. Difficulty, explanation style and next action all carry a stated reason the teacher can read.
  • Retrieval you can audit. Every tutor turn stores the chunk ids it used, and the dashboard shows them next to the reply.
  • Memory that resumes exactly. A student who closes the tab mid-question returns to the same phase and step, and a student who never returns still produces a partial learning trace.
  • A student never sees a note about themselves, enforced in three places: rules, API projection, and the model's own output gate.
  • 125 commits in three days, solo, shipping a Go API, a Three.js game, a realtime dashboard and a 3D asset pipeline to production on Cloud Run.

What we learned

The strongest tutoring agent is not the one that generates the most content; it is the one that turns small signals — a chosen reason, a "confusing" tap, a repeated slip, a tab closed mid-question — into a decision it can explain and a memory it can act on later.

Retrieval did not need embeddings to be useful. Chunks distilled from the teacher's own material, retrieved lexically and shown to the teacher by id, earned more trust than a black-box vector search would have in a classroom.

Provenance matters as much as the answer. Labelling every block as Gemini, fallback or example is what lets a teacher — or a judge — trust what they see.

And, on a personal note: with a 3D world eating every hour, the discipline that saved the project was asking of each feature, "does this make the agent's behaviour more visible?" and cutting everything that did not.

What's next for ClassGame

I did not build ClassGame to win a hackathon. I built it because I have watched teachers spend their evenings guessing what went wrong in thirty heads, and I believe that time belongs to their families and their students, not to a pile of paper. Winning would let me stop treating this as a three-day sprint and start treating it as the thing I actually want to spend the next years on: making a teacher's life lighter and a child's learning better, at the same time, with the same product.

Here is the dream, in the order I want to chase it.

More worlds, more subjects. Today Mia teaches Grade 5 math on one chalkboard. I want her in a science lab where a wrong hypothesis makes the beaker fizz, in a history classroom where the map redraws itself as the student argues, in a language room where the conversation partner is another chibi. Every scenario is a new set of misconceptions to model — and a new way for a kid to want to do the optional homework.

Play that teaches. The 3D room is still mostly a stage. I want the game itself to carry the pedagogy: the "pick the next move" equation duels I prototyped, cooperative missions where two classmates have to explain to each other before Mia lets them advance, a classroom that visibly changes as the class masters something together. Fun is not decoration here; it is the reason a student comes back on a Saturday.

Rewards that follow the thinking. Today a student earns the next stage and a progress bar that fills — honest, but thin. I want the loop to pay off: points for explaining a mistake rather than for guessing right, streaks for coming back on their own, characters and classroom items unlocked by mastering a topic instead of by grinding, and a class goal the group only reaches together. One rule holds it in place: the reward follows the reasoning. A student who answers wrong, says why, and gets there should out-earn one who guessed correctly — otherwise the game teaches the opposite of what Mia does.

A tutor that gets wiser. Semantic retrieval over the teacher's material next to the traceable lexical pass, memory that spans weeks and subjects instead of one mission, and Mia's voice on the chalkboard so a student who struggles with reading is not locked out of the help.

Scale without losing the notebook. The tenancy is already in every record — organization, district, school. I want a district to see, across hundreds of classrooms, which misconceptions travel together, and a new teacher to open ClassGame on day one and inherit what Mia already knows about her students. All of it with the same rule that shaped this project: the student never reads a note about themselves, and the teacher always sees the evidence.

Real classrooms. First pilots with the teachers whose complaint started this. Their corrections, not my roadmap, decide what ships next.

If ClassGame gets that chance, the win is not the prize. It is a teacher opening the dashboard on Monday morning, already knowing why — and a student walking up to a glowing desk because they want to.

Built With

+ 20 more
Share this project:

Updates

Submission history