Inspiration

Studying from long PDF and lecture documents often means spending a lot of time creating practice questions manually. We wanted to build a system that could turn existing study materials into interactive tests automatically. This idea led us to build QuizForge, an AI-powered educational assessment generator.

What We Built

QuizForge accepts PDF, DOCX, and PPTX documents, extracts and cleans the educational content, divides it into meaningful chunks, and uses large language models to generate multiple-choice questions. The generated questions are structured as JSON and passed through a validation layer to ensure that each question has four unique options, a valid answer, an explanation, and an appropriate difficulty level.

The application provides an interactive Streamlit quiz where users can answer the generated questions and receive their score, percentage, correct and incorrect answers, explanations, and difficulty information.

We experimented with two different LLM approaches: a locally deployed Llama 3.2-3B-Instruct model and the Gemini API. The local Llama pipeline allowed us to explore model quantization and efficient local inference, while Gemini provided a cloud-based comparison point.

How We Built It

We developed the project in Python using PyMuPDF for PDF processing, python-docx for DOCX processing, and python-pptx for PowerPoint extraction. Hugging Face Transformers and a 4-bit quantized Llama model were used for local LLM inference, while the Gemini API was integrated as an alternative generation approach. Streamlit was used to build the interactive interface.

A major focus was making the LLM output reliable. Since language models can occasionally produce malformed JSON or inconsistent answers, we implemented JSON recovery and MCQ validation rather than assuming every model response would be correct.

Challenges and What We Learned

One of the biggest challenges was connecting unpredictable LLM output to a deterministic application. Natural-language responses are difficult for programs to use reliably, so we learned how important structured output, validation, error recovery, and careful prompt design are when building LLM applications.

We also learned how document processing, chunking, model inference, and application logic need to work together as a single pipeline. Working with a locally quantized Llama model gave us practical experience with GPU-based inference and memory constraints, while integrating Gemini helped us understand the trade-offs between local and API-based LLM systems.

QuizForge ultimately became more than a document-to-question generator: it gave us hands-on experience building an end-to-end AI application, from raw educational documents to a usable interactive assessment system.

Built With

Share this project:

Updates

Submission history