Inspiration
Research papers contain valuable ideas, but they are often difficult to understand without strong background knowledge, significant time, and repeated reading. A 15-page paper can take hours to break down because the reader must understand the research problem, architecture, equations, experimental setup, and results at the same time.
I was inspired by visual educators such as 3Blue1Brown, who make difficult mathematical concepts intuitive through animation, and beginner-friendly technical educators who explain ideas conversationally instead of simply repeating definitions.
This led to PaperFlow31, an AI-powered learning platform that turns a research paper into a structured visual learning experience rather than another text summary.
What it does
PaperFlow31 allows a learner to upload a PDF research paper and automatically creates an interactive learning journey containing:
- A clear overview of the paper and its main contribution
- A structured four-chapter learning path
- A narrated visual lesson
- Interactive concept and signal explorations
- Paper-grounded questions and answers powered by Qwen
- Five review flashcards
- A five-question knowledge quiz
- Chapter-level explanations and visual storyboards
The experience is organized into five learner-focused stages:
- Overview – Understand the research objective, contribution, and learning path.
- Watch – Follow a narrated visual explanation of the paper.
- Explore – Interact with signals, model behavior, and important concepts.
- Quiz – Test understanding with immediate feedback.
- Review – Reinforce the most important ideas using flashcards.
For the CT-Net fNIRS research paper, PaperFlow31 provides a deeply designed visual lesson that explains:
- Why brain signals are difficult to classify
- How fNIRS measures changes in oxygenated and deoxygenated blood
- How CNNs capture local temporal patterns
- How Transformers model longer-range relationships
- How CT-Net combines both approaches
- How Grad-CAM helps explain the model’s predictions
The built-in Ask Qwen assistant is available throughout the lesson. It receives the uploaded paper text, the current page context, the active chapter, and relevant learning content so that its answers remain grounded in the source paper.
How we built it
PaperFlow31 uses a full-stack architecture designed for document processing, AI generation, interactive learning, and visual rendering.
AI and content generation
The backend extracts text from uploaded PDFs and sends carefully structured prompts to Qwen.
Qwen generates:
- A concise paper summary
- Four lesson chapters
- Learning goals for each chapter
- Flashcards
- Quiz questions
- Narration scripts
- Visual storyboard beats
- Context-aware answers to learner questions
The storyboard system uses a controlled visual schema with supported visual types such as:
signal_plotequation_transformarchitecture_flowattention_mapresult_comparisonconcept_diagram
This structure makes Qwen’s output easier to validate and safer to convert into visual experiences.
Visual lesson generation
The video pipeline uses Manim to create mathematical and scientific animations.
For the showcase CT-Net lesson, I created a continuous narrated animation containing more than 30 coordinated visual transitions. It moves from a noisy-signal analogy to fNIRS sensing, physiological interference, CNN feature extraction, Transformer attention, model prediction, and explainability.
The video also includes synchronized narration, H.264 video, AAC audio, browser range-request support, and cached delivery for fast playback.
Backend
The backend was built with:
- Python
- FastAPI
- Qwen API
- PyPDF
- Manim
- gTTS
- FFmpeg
- Pydantic
The API handles:
- PDF uploads
- Text extraction
- Lesson generation
- Storyboard generation
- Question answering
- Video rendering
- Video status checks
- Cached video delivery
- HTTP range requests for browser playback
Generated lessons and storyboards are cached to prevent unnecessary repeated Qwen calls and improve navigation performance.
Frontend
The learner experience was built with:
- Next.js
- React
- TypeScript
- Tailwind CSS
- Plotly
The interface was designed as a polished learning product rather than a basic document-processing dashboard. It includes responsive navigation, premium visual hierarchy, progress indicators, interactive quizzes, animated flashcards, contextual AI assistance, and dedicated learning modes.
Deployment
The entire project is containerized using Docker.
The current public deployment uses:
- Vercel for the Next.js frontend
- Render for the FastAPI backend
- GitHub for source control and automated deployments
The Docker architecture also makes the application portable to Alibaba Cloud ECS after account verification is completed.
Challenges we ran into
One major challenge was converting AI-generated explanations into useful visual scenes. Free-form model output cannot be executed directly as animation code because it may be incomplete, inconsistent, or unsafe. I addressed this by introducing a structured storyboard format with a limited set of supported visual types and validated configuration objects.
Another challenge was synchronizing narration with continuous animation. A useful educational video requires more than showing text beside audio. Every concept must appear at the right time, remain visible long enough to understand, and transition naturally into the next idea.
Container networking was another important challenge. Browser-side requests, Next.js server-side requests, and Docker-internal requests require different API addresses. I separated the public API URL from the internal server API URL so the same application works locally, inside Docker, and in production.
Production performance also required additional work. Initially, page navigation repeatedly regenerated lesson content and storyboards. I added caching to both the backend and Next.js data layer and disabled unnecessary eager prefetching.
Finally, free cloud services can enter sleep mode after inactivity. The application now reduces repeated AI calls and reuses generated content, but the first request after a long idle period may still take longer while the backend wakes up.
Accomplishments that we're proud of
I am proud that PaperFlow31 goes beyond summarizing a PDF.
The project delivers an end-to-end learning experience that combines:
- Research-paper understanding
- Structured Qwen generation
- Visual storytelling
- Scientific animation
- Interactive exploration
- Knowledge assessment
- Context-aware question answering
The CT-Net showcase lesson contains a polished narrated animation with more than 30 visual transitions and explains a complex CNN-Transformer architecture in an intuitive sequence.
I am also proud that the application is fully deployed, Dockerized, publicly accessible, and designed with a clear production architecture rather than being limited to a local prototype.
What we learned
The most important lesson was that generating educational content is different from generating a summary.
A strong learning experience requires several layers:
- Identifying the learner’s likely knowledge gap
- Selecting the correct concepts
- Ordering those concepts carefully
- Choosing the right visual representation
- Connecting narration to animation
- Testing understanding
- Giving learners a way to ask follow-up questions
I also learned that LLMs are most effective when they are given clearly defined responsibilities. Qwen performs well when generating explanations, lesson structures, storyboards, and grounded answers, while deterministic application code handles validation, rendering, caching, navigation, and safety.
The project also reinforced the importance of production concerns such as container networking, persistent storage, CORS, video streaming, deployment portability, and latency optimization.
What's next for PaperFlow31
The next phase is to make the visual-generation pipeline more general across different research domains.
Planned improvements include:
- A validated animation DSL generated by Qwen
- Automatic equation extraction and transformation
- Diagram generation from architecture descriptions
- Figure and table extraction from PDFs
- Citation links from every explanation back to the original page
- Persistent user libraries and learning history
- Adaptive quizzes based on weak concepts
- Multiple explanation levels for beginners, students, and researchers
- Support for additional languages and narration voices
- Collaborative lessons for classrooms and research groups
- Deployment on Alibaba Cloud ECS with persistent storage
- Alibaba Cloud Object Storage Service for uploaded papers and generated videos
- Background rendering queues for scalable visual generation
The long-term goal is for PaperFlow31 to become an AI learning engine that can transform difficult research into an experience people can watch, explore, question, and remember.
Log in or sign up for Devpost to join the conversation.