Inspiration
Most of us have sat through a confusing explanation and only gotten it once someone drew it out or explained it differently. We kept seeing AI tutors that just gave more text, and it felt like a missed opportunity for the majority of students who don't learn that way.
What It Does
HoosLearn is an AI-driven platform that turns any topic into an immersive learning experience. Instead of returning a wall of text, users choose how they want to learn:
- Songs: educational lyrics that make concepts actually stick
- Videos: short narrated clips that walk you through the material
The goal is simple: make learning more engaging, memorable, and tailored to how individual students actually think.
How We Built It
HoosLearn runs on a multi-model pipeline across three layers:
| Layer | What It Does |
|---|---|
| Input | User enters a topic and selects a learning format |
| Generation | Routes to the right model. ElevenLabs for music & video, Claude for scripting |
| Delivery | Renders output in real time + generates a persistent shareable link |
Gemini handles prompt translation, converting raw topics into format-specific scripts that genuinely teach rather than just describe.
Challenges We Ran Into
- Latency: generating rich media in real time is slow. We optimized through parallel requests and smart pre-processing, but it remained our biggest constraint
- Consistency across formats: getting a song and a video to accurately represent the same concept without one going off the rails required careful per-modality prompt engineering
- Prompt translation: converting a single topic into different media formats without losing meaning was harder than expected
- Avoiding generic output: AI content has a recognizable sameness to it. We pushed hard to make results feel educational, not just technically correct
Accomplishments We're Proud Of
- Built a fully functional multimodal learning platform in 24 hours
- Unified multiple AI models into one coherent pipeline with consistent output
- Reduced the user journey to just a few clicks
- Made the song format genuinely educational, not just a novelty
- Shipped persistent shareable links so experiences can be revisited and passed around
What We Learned
- Multimodal learning works: audio and video improve retention in ways text alone doesn't
- Prompting is format-specific: what works for text fails for audio, what works for audio fails for video
- Latency is a product problem: slow generation breaks user experience regardless of output quality
- The details matter: clarity, speed, and simplicity had more impact on how the product felt than any individual feature
What's Next for HoosLearn
- Personalized learning paths: adapt outputs over time based on how individual users learn best
- Interactive experiences: quizzes, simulations, and adaptive feedback beyond passive content
- Deeper AI orchestration: agentic workflows that refine and validate outputs before delivery
- Collaboration tools: share, remix, and build on each other's generated content
- Classroom integrations: teacher-facing tools that bring HoosLearn into structured learning environments
Built With
- amazon-web-services
- elevenlabs
- fastapi
- ffmpeg
- gemini
- react
- redis
- typescript
- websockets
Log in or sign up for Devpost to join the conversation.