NeuralFlix โ€” Devpost Project Story

Paste this directly into the "About the project" field on Devpost.
Markdown renders natively. LaTeX math is supported.


The Project Story (copy everything below this line)


๐ŸŽฌ Inspiration

Every major streaming platform โ€” Netflix, Prime, Disney+ โ€” is built around the same bias: Hollywood first, everything else as an afterthought.

I grew up watching Bollywood, Korean thrillers, Iranian art-house films, and Japanese animation alongside English blockbusters. When I started using recommendation engines seriously, I noticed something frustrating: none of them understood that a person who loves Parasite also probably wants to discover Asghar Farhadi, or that a fan of Interstellar might connect deeply with Arrival and Tumbbad.

NeuralFlix was born from one question: What would a recommendation engine look like if it was designed for a truly global audience โ€” not just an English-speaking one?

I also wanted to prove something to myself: that you could build a production-grade, multi-model ML recommendation system from scratch โ€” not just wrap an API.


๐Ÿง  What I Built

NeuralFlix is a full-stack ML-powered movie recommendation platform with a global cinema catalog spanning 10,000+ curated films across 12 cinema regions โ€” Bollywood, Hollywood, Korean, Japanese, French, Spanish, Tollywood, Tamil, Chinese, Iranian, Turkish, and Portuguese cinema.

The platform is end-to-end: from data ingestion to ML training to a premium animated web interface โ€” deployed live on Vercel + Render.

The ML Stack (7 Models Working Together)

The core of NeuralFlix is RecommenderOS โ€” a multi-stage hybrid recommendation pipeline I designed from scratch:

Stage 1 โ€” Recall (parallel retrieval across 4 models):

1. Neural Collaborative Filtering (NCF)
A dual-stream PyTorch neural network combining:

  • GMF (Generalized Matrix Factorization) โ€” models linear latent correlations between users and items
  • MLP โ€” captures complex, non-linear higher-order interaction patterns
  • NeuMF โ€” fuses GMF and MLP outputs for precise interaction prediction

2. Sequential Transformer (SASRec)
A self-attention transformer that treats a user's watch history as a sentence โ€” each film is a token. Instead of assuming static taste, it weights recent interactions more heavily and predicts the next movie the user is likely to enjoy. The math behind it:

$$\text{Attention}(Q, K, V) = \text{softmax}!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V$$

3. Content-Based TF-IDF Engine
Builds a 384-dimensional TF-IDF matrix across movie overviews, genres, director names, and cast โ€” enabling semantic similarity retrieval. When you click a movie, the content engine finds structurally similar films across all 12 cinema regions.

4. Graph Neural Network (GNN)
Models movies, actors, directors, genres, and users as nodes in a heterogeneous bipartite graph. Edge predictions surface deep multi-hop connections โ€” finding that a fan of a director's early work also connects to a foreign film with the same cinematographer.

Stage 2 โ€” Blend:

$$\text{FinalScore} = \Big( \alpha \cdot \text{NCF_Score} + (1-\alpha) \cdot \text{Semantic_Similarity} \Big) \cdot \text{RegionalBoost}$$

where \(\alpha\) dynamically adjusts based on how much interaction data exists for the user, and \(\text{RegionalBoost}\) amplifies results from cinema regions matching the user's taste profile.

Stage 3 โ€” Re-rank, Diversify, Explore:

  • LightGBM Ranker โ€” re-scores candidates using feature vectors: genre match, rating, decade match, votes, normalized popularity
  • K-Means Diversity Layer โ€” clusters candidates by genre embeddings; ensures maximum 2 results per genre in the final top-20
  • Thompson Sampling Bandit โ€” 15% exploration chance: swaps one result with a high-quality film outside the user's comfort zone, expanding taste over time
  • Sentiment Re-ranker โ€” adjusts scores based on aggregated review sentiment from IMDb/Trakt

๐Ÿ—„๏ธ Data Architecture

The platform ingests and serves from a 3-database architecture:

Database Role
MongoDB Atlas 10,000+ movie documents with full metadata, genres, region tags, popularity scores, embeddings
PostgreSQL Telemetry logs, user interaction events, A/B test signals
Redis Recommendation cache (10-min TTL), session tokens, rate limiting

External APIs integrated:

  • TMDB API โ€” primary metadata, cast/crew, poster art, trailers
  • OMDb / IMDb โ€” ratings, awards, box office data
  • Trakt.tv โ€” community trending signals, real watchlist data
  • Watchmode API โ€” streaming availability (Netflix, Prime, Hotstar, ZEE5, Crunchyroll, MUBI, etc.)

The catalog distribution is deliberately balanced โ€” combating the Hollywood-bias problem:

Region Target % Count
Hollywood (English) 30% ~3,000
Bollywood (Hindi) 12% ~1,200
Korean 10% ~1,000
Japanese 8% ~800
French 6% ~600
Spanish 5% ~500
Tollywood/Tamil/Others 29% ~2,900

๐ŸŽจ Frontend Architecture

The frontend is Next.js 15 with App Router, styled with Tailwind CSS v4 and a custom design token system. Key UI features:

  • Premium Light/Dark Theme โ€” custom CSS token system with warm cream light mode (#FAFAF7 surface) and deep obsidian dark mode (#06060A)
  • Framer Motion animations โ€” page transitions, scroll reveals, staggered card grids, animated counter stats
  • Three.js WebGL layer โ€” background particle canvas, 3D movie card tilt on hover, interactive recommendation orb
  • Interactive Cinema World Map โ€” SVG globe with clickable countries navigating to regional cinema catalogs
  • Taste DNA Visualizer โ€” radar chart mapping a user's genre affinity, preferred decades, top directors, language taste
  • Mood Discovery Engine โ€” emotional sliders translate to vector search queries
  • Command Palette (Cmd+K) โ€” spotlight-style unified search, navigation, and system commands
  • Real-time WebSocket โ€” recommendation feed updates as user browses

๐Ÿ”ง How I Built It

Backend (FastAPI)

backend/
โ”œโ”€โ”€ main.py              # Lifespan, CORS, middleware, router registration
โ”œโ”€โ”€ ml/                  # 12 ML model files
โ”‚   โ”œโ”€โ”€ ncf_model.py     # PyTorch NCF (GMF + MLP + NeuMF)
โ”‚   โ”œโ”€โ”€ sasrec_model.py  # Self-attention sequential transformer
โ”‚   โ”œโ”€โ”€ gnn_model.py     # Bipartite graph neural network
โ”‚   โ”œโ”€โ”€ content_based.py # TF-IDF semantic similarity
โ”‚   โ”œโ”€โ”€ hybrid_recommender.py  # Pipeline composer
โ”‚   โ”œโ”€โ”€ ranker.py        # LightGBM candidate re-ranker
โ”‚   โ”œโ”€โ”€ diversity.py     # K-Means diversity injection
โ”‚   โ”œโ”€โ”€ exploration_bandit.py  # Thompson sampling
โ”‚   โ”œโ”€โ”€ sentiment_reranker.py  # Review-weighted re-scorer
โ”‚   โ”œโ”€โ”€ cold_start.py    # Onboarding bootstrap
โ”‚   โ””โ”€โ”€ taste_profile.py # Taste DNA matrix builder
โ”œโ”€โ”€ routes/              # 12 API route files
โ”œโ”€โ”€ utils/               # TMDB, Trakt, Watchmode clients
โ””โ”€โ”€ api/websocket.py     # Real-time WS feed

Frontend (Next.js 15)

frontend-next/
โ”œโ”€โ”€ app/                 # 15 pages (App Router)
โ”‚   โ”œโ”€โ”€ page.tsx         # Homepage with regional rows
โ”‚   โ”œโ”€โ”€ discover/        # Full catalog browser with filters
โ”‚   โ”œโ”€โ”€ movie/[id]/      # Detail page: trailer, cast, similar
โ”‚   โ”œโ”€โ”€ recommendations/ # ML recommendation hub + Taste DNA
โ”‚   โ”œโ”€โ”€ cinema/[region]/ # Regional cinema explorers
โ”‚   โ”œโ”€โ”€ search/          # Full-text + mood search
โ”‚   โ”œโ”€โ”€ profile/         # Watch history, watchlist, ratings
โ”‚   โ”œโ”€โ”€ onboarding/      # Cold-start genre/movie selection
โ”‚   โ””โ”€โ”€ world-map/       # Interactive globe
โ”œโ”€โ”€ components/          # 30+ components
โ”‚   โ”œโ”€โ”€ three/           # Three.js WebGL components
โ”‚   โ”œโ”€โ”€ movie/           # Rating panels, OTT badges
โ”‚   โ””โ”€โ”€ TasteDNA.tsx     # Taste radar chart
โ””โ”€โ”€ styles/
    โ””โ”€โ”€ tokens.css       # 120+ CSS design tokens

Infrastructure

  • Render โ€” FastAPI backend, auto-deploy from GitHub
  • MongoDB Atlas โ€” M0 free tier for movie catalog
  • Redis (Render managed) โ€” caching and rate limiting
  • Vercel โ€” Next.js frontend, edge CDN
  • Docker Compose โ€” local full-stack development

๐Ÿ˜ค Challenges I Faced

1. The cold-start problem is genuinely hard.
When a new user registers, there's zero interaction data. I solved this with a 4-step onboarding flow: genre selection โ†’ curated movie picks โ†’ language preferences โ†’ initial taste profile generation. The cold-start engine then uses content-based similarity on the onboarding picks to generate the first recommendation set before NCF has any data to work with.

2. Balancing 7 ML models without any one dominating.
Each model has different scoring scales. NCF returns softmax probabilities (0โ€“1). Content-based returns cosine similarities (โ€“1 to 1). The ranker returns raw LightGBM scores. Getting all of these onto a comparable scale required careful normalization before blending โ€” I spent more time on this than on any single model.

3. Scale mismatch between development and production.
I initially hardcoded NCFModel(num_users=1000, num_items=5000) in development. When I started populating the real MongoDB catalog with 10,000 movies, the NCF was silently ignoring 50% of them โ€” because items with index > 5,000 didn't exist in the embedding matrix. Debugging this required tracing through the candidate pool โ†’ ID mapper โ†’ model index translation layer I subsequently built.

4. Geography bias in TMDB's own API.
TMDB's /discover/movie?sort_by=popularity.desc returns 90% English-language films on page 1. To achieve my target regional distribution, I built a custom ingestion scheduler that runs separate discovery queries per language code, enforces page limits per region, and fills gaps in underrepresented regions on a 12-hour schedule.

5. Making recommendations feel fast.
The full ML pipeline (NCF โ†’ SASRec โ†’ content โ†’ blend โ†’ rank โ†’ diversity โ†’ explain) takes 2โ€“4 seconds on first run. I solved this with Redis caching (10-min TTL per user), background SVD retraining (hourly, not per-request), and a lazy-loading pattern for the SentenceTransformer model that reduced cold startup from 60 seconds to 3 seconds.


๐Ÿ“š What I Learned

  • ML systems are mostly plumbing. The actual neural network code is maybe 20% of the work. The remaining 80% is data normalization, ID mapping, candidate sampling, cache invalidation, fallback chains, and graceful degradation when models aren't loaded.

  • Regional cinema deserves its own recommendation research. Most academic recommendation papers use MovieLens โ€” which is 95% Hollywood. Building a system that genuinely understands why a Bollywood fan might connect with a Korean melodrama required thinking beyond genre tags into cultural and emotional dimensions.

  • The diversity-accuracy tradeoff is real. A purely accuracy-optimized system creates a filter bubble fast. My early experiments showed users getting the same 5 genres repeatedly after 20 ratings. The K-Means diversity layer + Thompson Sampling bandit meaningfully improved catalog coverage in testing.

  • Full-stack ML requires a different mental model than either pure ML or pure web dev. The data schemas, API contracts, and UI states all have to be designed together โ€” not sequentially. I redesigned the movie data schema 3 times as I discovered what the ML pipeline needed versus what the frontend needed.


๐Ÿ—๏ธ What's Next

  • Qdrant vector search integration โ€” fully wired semantic mood search using dense embeddings
  • MovieLens 25M pre-training โ€” pre-train NCF on 25M real ratings before any NeuralFlix user interaction
  • Native mobile app โ€” React Native with offline recommendation cache
  • Collaborative playlists โ€” shared watchlists with ML-powered co-recommendation
  • Director/actor graph explorer โ€” visual network graph of cinematic connections

๐Ÿ› ๏ธ Built With

Python ยท FastAPI ยท PyTorch ยท scikit-learn ยท LightGBM ยท Surprise (SVD) ยท sentence-transformers ยท Next.js 15 ยท React 19 ยท TypeScript ยท Tailwind CSS v4 ยท Framer Motion ยท Three.js ยท MongoDB Atlas ยท PostgreSQL ยท Redis ยท Docker ยท Celery ยท TMDB API ยท Trakt.tv API ยท Watchmode API ยท OMDb API ยท Qdrant ยท Vercel ยท Render ยท Prometheus ยท Grafana ยท Structlog

Built With

Share this project:

Updates

Submission history