Inspiration
I wanted my first real machine learning project to be more than just a notebook full of accuracy scores — I wanted something anyone could actually open and use. Movie reviews felt like the perfect playground: everyone has opinions about films, the language is rich and emotional, and there's a genuinely useful question hiding in there — can a machine tell whether someone loved or hated something just from how they wrote about it? That question turned into the Sentiment Analyzer.
What it does
Sentiment Analyzer is a web dashboard that reads a movie review and instantly tells you whether it's positive or negative, along with a confidence score. Beyond live predictions, it also lets you:
- See exactly how well the underlying model performs (accuracy, precision, recall, confusion matrix) on data it's never seen before
- Explore word clouds that reveal which words are most strongly tied to positive vs. negative sentiment
- Browse real sample reviews from the training dataset, filtered by sentiment
It's built on the IMDB Dataset of 50,000 movie reviews and currently runs at ~88.6% accuracy on unseen test data.
How I built it
The pipeline started with cleaning 50,000 raw reviews — stripping leftover HTML tags from web scraping, lowercasing everything, removing punctuation, and dropping 418 exact-duplicate reviews that would otherwise leak between training and test data.
From there, I converted the cleaned text into numbers using TF-IDF (Term Frequency–Inverse Document Frequency), capped at the 5,000 most informative words, with English stop words removed automatically. I trained a Logistic Regression classifier on an 80/20 stratified split — chosen deliberately because linear models like this perform strongly and efficiently on the kind of high-dimensional, sparse features TF-IDF produces, without needing a GPU or long training times.
Once the model was trained and evaluated, I serialized it with joblib and wrapped it in a Streamlit dashboard with custom theming, an animated splash screen, and four interactive pages. The whole thing is version-controlled on GitHub and deployed live on Streamlit Community Cloud.
Challenges I ran into
The trickiest bug appeared after deployment: my live "Model Stats" page suddenly reported ~90.9% accuracy — noticeably higher than the 88.6% from my original training notebook. At first this looked like good news, but it didn't add up.
Digging in, I found two mismatches between the notebook and the deployed app: the app's text-cleaning function was slightly different from the one used during training, and — more importantly — the app was re-splitting the full, un-deduplicated dataset for evaluation. That meant duplicate reviews could end up in both the "training" and "test" portions of that internal split, letting the model effectively see answers it had already memorized. Once I aligned the deployed app's preprocessing exactly with the notebook's, the numbers became consistent and trustworthy again.
Accomplishments that I'm proud of
I'm proud that this isn't just a model sitting in a notebook — it's a fully deployed, publicly accessible product that anyone can try right now. I'm also proud of catching and correctly diagnosing the duplicate-data-leakage bug myself; recognizing that an unexpectedly better number is still a bug worth investigating, not just a lucky result, felt like a genuine engineering instinct clicking into place.
What I learned
Beyond the mechanics of TF-IDF and Logistic Regression, the biggest lesson was that a model is only as trustworthy as the pipeline serving it. Small inconsistencies between how you train a model and how you deploy it can silently distort your results — and the discipline of keeping preprocessing identical across environments matters just as much as the choice of algorithm itself. I also learned firsthand why stratified splitting and deduplication aren't just best-practice checkboxes — skipping them has real, measurable consequences.
What's next for Sentiment Analyzer
- Experiment with a transformer-based model (e.g. DistilBERT) to see how much accuracy improves on the negation and sarcasm cases that trip up a bag-of-words approach like TF-IDF
- Extend beyond movie reviews to other domains (product reviews, restaurant reviews) and test how well the model generalizes without retraining
- Add batch prediction support, so users can upload a CSV of reviews and get sentiment predictions for all of them at once
Log in or sign up for Devpost to join the conversation.