Inspiration

What it does

How we built it

Challenges we ran into

BAU EduHub — Past Papers, Powered by AI

Inspiration

Every semester at Bangladesh Agricultural University, the same scramble happens right before exams: students hunting down past question papers through scattered Facebook groups, WhatsApp threads, or whichever senior happens to still have a folder from three years ago. There's no central place to find them, no way to know if a paper is even the right version, and definitely no way to actually learn from the papers beyond just reading through them once.

We kept asking ourselves: why does something so fundamental to exam prep still run on informal networks and luck? Students who don't have the right connections are simply at a disadvantage — not because they study less, but because they have less access. That gap is what pushed us to build something that treats past questions not as scattered files, but as a shared institutional resource — and then layers AI on top to make that resource genuinely useful, not just searchable.

What it does

BAU EduHub lets students upload past exam question papers after each exam, tagged automatically by course, semester, and topic using AI-based extraction. On the other end, any student can search, filter, and download these papers instantly instead of digging through group chats.

But the real value goes beyond storage. Once enough papers accumulate, our AI assistant analyzes them as a dataset — clustering repeated topics across years, highlighting high-frequency questions, and even generating mock quizzes built directly from real past questions. It turns a passive PDF archive into an active study tool that gets smarter the more the community contributes.

How we built it

We split the system into modules that could be built and tested independently, which mattered a lot given our time constraints:

  • User module for authentication and department/batch profiles, so content stays relevant to each student
  • Upload & ingestion pipeline that runs uploaded papers through OCR and an LLM to auto-tag course, topic, and question type
  • Question bank module for search, filtering, verification, and download
  • AI assistant module built on a retrieval-augmented generation (RAG) approach — the model answers questions and generates quizzes using the actual uploaded papers as its knowledge base, rather than relying on general knowledge alone
  • Admin module for moderation and duplicate detection

On the technical side, we used a vector-based similarity search to power both duplicate detection and semantic search, so a query for "nutrient cycle" surfaces relevant questions even if they're phrased differently across years. If we denote the embedding of an uploaded question as $q_i$ and an existing question as $q_j$, we flag likely duplicates or close matches using cosine similarity:

$$ \text{sim}(q_i, q_j) = \frac{q_i \cdot q_j}{\lVert q_i \rVert \, \lVert q_j \rVert} $$

Pairs above a similarity threshold get grouped for topic clustering instead of treated as separate questions — this is also what lets the AI assistant identify "commonly asked" topics across years.

Challenges we ran into

  • OCR accuracy on handwritten or scanned papers — a lot of real question papers are photographed, not typed, so getting clean text extraction before tagging took real tuning.
  • Duplicate detection at scale — naive exact-match checking missed near-duplicate questions with slightly different wording, which is why we moved to embedding-based similarity instead.
  • Balancing automation with trust — we didn't want AI auto-tagging to silently mislabel a paper, so we added a lightweight verification/upvote system so the community can flag and correct errors rather than relying purely on the model.
  • Scoping for hackathon time — the idea naturally wanted to grow into a full LMS, so we had to consciously cut features (like full course chatbots for every subject) down to a working core: upload, tag, search, and AI-generated quizzes from real data.

What we learned

We came away with a much better sense of how RAG systems behave when the "knowledge base" is community-contributed rather than curated — messy, inconsistent, and constantly growing, which is a very different problem from querying a clean fixed document set. We also learned how much of a "smart" AI feature actually depends on unglamorous groundwork: OCR quality, tagging consistency, and duplicate handling mattered more to the final experience than the LLM prompt itself.

What's next for BAU EduHub

  • Full RAG-based course chatbots for every department, not just the question bank
  • Personalized mock tests based on a student's weak topics, tracked over time
  • Expanding beyond BAU to other public universities in Bangladesh, where the same access gap exists ## Accomplishments that we're proud of

What we learned

What's next for Smart Exam Preparation & Previous Question Repository

Built With

Share this project:

Updates

Submission history