Inspiration

Most AI study tools work in one direction: you ask, they explain, and you leave feeling like you understand. That feeling is easy to fake, because reading a clear explanation isn't the same as being able to give one. The classic way to check is to teach the idea to someone else, but a student cramming alone at night doesn't have a classmate to teach. We wanted to build that classmate, one who has only seen this week's lectures, gets confused in realistic ways, and tells you honestly where your explanation fell apart.

What it does

Pupil flips the usual AI tutor around: you teach, the AI learns.

A student picks a subject and one week, a range, or any combination of weeks, then ticks which lecture files the AI student is allowed to know. They choose a persona: Pip, a curious kid who needs simple words; Sage, a sceptic who wants examples and failure cases; or Milo, who holds a wrong belief and only lets go when you explain why it's wrong. They teach by voice or text for 10, 15 or 25 minutes. Each idea gets up to three attempts, with a hint button if they're stuck. Every answer is graded as explained, shaky or wrong, and judged against the student's own slides, lab notebooks or recording. At the end they get an Understanding Map of the week's ideas, coloured by how well they explained each one. Reteach restarts on just the shaky ideas and gaps.

Around the teaching mode, Ask answers questions about deadlines and lectures with citations to the slide, page or recording. Brainstorm writes quizzes from the student's own slides.

How we built it

Frontend: Angular 21 with pages for login, home, setup, the live session, the map and Ask. Backend: FastAPI, with SQLite locally or Postgres with pgvector. Canvas sync: it uses the student's own Canvas token, stored encrypted and never sent back to the browser, and only documented Canvas endpoints. It ingests slide PDFs, lab notebooks and lecture recordings, which are transcribed with Whisper. Search: content is chunked and embedded locally with fastembed. Concept graph: it links ideas across weeks with relations like "builds on" and "part of". The graph decides which ideas to teach and how the map is laid out, and the vector search supplies the original lecture evidence. Scoping: every lookup is filtered by student, course, weeks and ticked files, so the AI student can't know more than the student chose to share. LLM layer: any OpenAI-compatible model, including a local Ollama server. Replies and grading use structured JSON output. Quality checks: more than 90 automated tests, plus an evaluation harness for retrieval and answers.

Challenges we ran into

Making the model act like a student, not a tutor. Models kept swapping roles, greeting their own persona or starting to explain. We added checks to catch that, plus a small-talk detector so a "hi, can you hear me?" doesn't get graded as a bad explanation. Keeping the AI's knowledge limited. "Only knows what you ticked" needed filtering through the vector search and the concept graph, so an idea is only taught if its evidence is in the ticked files. Reliability with local models. Constrained JSON generation can loop, so we capped output length and timeouts. If the LLM fails or isn't configured, a key-term check keeps sessions working. Data we can't get. There's no approved lecture-capture API available to us, so recordings come in by manual upload instead of being scraped. Privacy. Student Canvas data is sensitive, so encryption, per-course authorisation and tokens that never come from the request body had to be there from the start.

Accomplishments that we're proud of

We went from a front-end mock to a working end-to-end system: real Canvas content in, graded teaching sessions out. It runs fully on a laptop with a local model, with no cloud key needed. It degrades gracefully: with no LLM, it still works. Every grade and answer points back to a specific slide, page or timestamp, so students can check it. More than 90 tests cover the security and isolation rules, so one student's data can't leak into another's.

What we learned

Teaching exposes gaps that reading never does. The hardest part of building this was the same as the hardest part of using it. Grounding matters more than model size. Retrieval filtered to the right material did more for trust than any prompt tweak. Each persona needs its own structure. A prompt alone wasn't enough to keep "the mixed-up one" mixed up in the right way. Fallbacks are a feature. Building a no-LLM path first made the system easier to test and more dependable. For education tools, privacy and scope aren't extras to add later.

What's next for Pupil

Lecture capture: build the connector for an approved integration, with the university's cooperation, so recordings flow in automatically. Better voice: stronger speech handling for talking through ideas naturally. Spaced repetition: resurface weak ideas over the semester, not just in the next session. More tools for students and tutors: share a map with a tutor and save sessions as revision notes. Wider support: more courses, more file types, and tuning of the personas and grading using real student feedback.

Built With

+ 7 more
Share this project:

Updates

Submission history