Inspiration As students tackling dense subjects like algorithms, data structures, and microeconomics, we realized that traditional study methods are broken. Most AI summarization tools just compress text into bullet points, encouraging rote memorization of disjointed facts. But passing rigorous exams requires understanding the underlying theory, not just the surface-level definitions. We wanted to build a tool that forces active recall and tests true comprehension—like knowing exactly why an O(n log n) sorting algorithm behaves the way it does, or understanding the precise mechanism of a Left-Right AVL tree rotation.

What it does Anki Pipeline is an automated, cognitive-science-backed study tool. Users upload raw, messy lecture slides (PDFs) or paste dense text. The pipeline automatically cleans the text, extracts the core theoretical concepts, and generates a set of highly difficult, academic multiple-choice questions. It specifically engineers "plausible distractors"—wrong answers based on common student misconceptions. Finally, it formats these questions into a clean JSON structure and instantly exports a CSV file ready to be imported directly into Anki for spaced repetition.

How we built it We built a decoupled architecture prioritizing speed and a clean user experience:

Frontend: A sleek, dark-mode web interface built with HTML/CSS/Vanilla JS, utilizing native browser APIs for dynamic CSV generation and downloading.

Backend: A lightweight, lightning-fast Python server powered by FastAPI and Uvicorn.

Data Parsing: We integrated PyMuPDF (Fitz) to seamlessly extract text from uploaded PDF slides.

AI Engine: We utilized structured, zero-shot prompt engineering to force the LLM into returning strict JSON schemas, allowing our backend to parse the data safely without regex hacks.

Challenges we ran into The midnight API battles! At around 12:00 AM, our multi-stage generation pipeline started hitting severe 429 Too Many Requests rate limits on the free tier. When we tried to migrate endpoints, we ran straight into 404 Not Found deprecation errors for older models. With the clock ticking, we had to rapidly refactor our backend from a multi-stage sequential loop into a "Single-Shot Pipeline" to minimize API calls. Finally, to ensure a flawless presentation for our demo, we engineered a robust local mock-engine fallback using asyncio to simulate network latency, guaranteeing our UI and CSV export logic could be demonstrated perfectly regardless of third-party server stability.

Accomplishments that we're proud of We are incredibly proud of the prompt engineering that generates the distractors. Standard AI quizzes are too easy; our pipeline generates wrong answers that actually make you think. We are also proud of building a fully functional, end-to-end pipeline (from raw PDF upload to formatted CSV download) in a single night without relying on heavy frontend frameworks.

What we learned We learned the hard way that third-party APIs are volatile and unpredictable. Building resilient software isn't just about writing the "happy path" code; it's about handling edge cases, managing asynchronous operations gracefully, and knowing when to pivot your architecture to meet a strict deadline.

What's next for Anki Pipeline

Production API Integration: Upgrading to a paid API tier to remove rate limits and allow for processing massive, 100-page slide decks.

Multimodal Parsing: Using vision models to extract formulas, graphs, and diagrams from the slides and embed them directly into the flashcards.

Native Spaced Repetition: Building our own review algorithm directly into the web app so users don't even need to export to Anki.

A Note on the Demo: Due to strict 'Rate Limit: 0' and 'Quota Exhausted' blocks on the Gemini 2.0 Flash API during the final hours of the build, we implemented a Service Virtualization layer. This means that while the PDF parsing logic is live, the final LLM response in this demo is served by a local mock engine. This ensures the UI stability and CSV export integrity required for a final presentation, while demonstrating exactly how the multi-stage pipeline handles structured JSON data once the API quota is restored.

Built With

Share this project:

Updates

Submission history