Inspiration
As a grad student, I spend hours after every lecture re-reading slides and PDFs, trying to figure out what actually matters enough to be tested on. Generic AI quiz tools exist, but they pull from the model's general knowledge, not from what my professor actually taught — so the questions either miss the point of the lecture or invent details that were never covered. I wanted something that only quizzes me on my own material, and tells me exactly where in my notes the answer came from.
What it does
StudyBuddy AI lets a student upload their own lecture slides (PPTX) or notes (PDF), pick a topic, and get a multiple-choice quiz generated only from that document. Every question comes with an explanation and a citation back to the exact slide or page it was grounded in, so a student can immediately go check the source instead of just trusting the AI.
How I built it
- Extraction —
pdfplumberandpython-pptxpull raw text out of uploaded PDFs and slide decks, page by page or slide by slide. - Chunking — slide decks are split one-slide-per-chunk, since a slide is already a single idea. PDFs use a sliding-window split with overlap, since prose doesn't have the same natural boundaries.
- Embedding + retrieval — each chunk is embedded with a
Sentence-Transformers model (
all-MiniLM-L6-v2) and indexed in FAISS using cosine similarity, so when a student types a topic, I retrieve the most relevant chunks of their own material. - Grounded generation — the retrieved chunks are passed to Gemini with an explicit instruction to generate questions only from the provided text, returning structured JSON (question, options, correct answer, explanation, source location) — not freeform text — so it can be rendered and graded reliably.
- UI — a Streamlit app ties it together: upload, pick a topic, take the quiz, get graded with per-question feedback and the source citation.
Challenges I ran into
- Keeping the AI grounded. The first versions of the prompt let Gemini drift into generating textbook-style questions that sounded right but weren't actually supported by the uploaded slides. Tightening the prompt to explicitly forbid outside knowledge, and forcing a citation field in the output schema, fixed this.
- Environment setup. A conflicting Python installation on my machine was silently creating broken virtual environments, which broke FAISS's build process through a chain of SSL errors. Diagnosing this down to the actual Python executable being used, rather than just the command name, was its own debugging exercise.
- Doing this solo. Building the full retrieval-and-generation pipeline, the Streamlit front end, and the grounding/citation logic alone, from scratch, in one weekend, meant being deliberate about what to cut — I kept the chunking and retrieval simple on purpose so the core idea (grounded, citable quizzes) stayed solid under time pressure.
Built With
- faiss
- gemini-api
- google-generativeai
- machine-learning
- natural-language-processing
- numpy
- pdfplumber
- python
- python-pptx
- rag
- sentence-transformers
- streamlit
Log in or sign up for Devpost to join the conversation.