💡 Inspiration

With the rapid growth of digital communication, phishing attacks have become more frequent and increasingly sophisticated. Many existing solutions focus only on blocking links, but modern scams exploit language—using urgency, fear, and manipulation—to deceive users.

This inspired us to build a Machine Learning–based solution that understands the language patterns of phishing messages and helps everyday users quickly identify whether a message is safe or a potential scam.


⚙️ What it does

Phishing Detector is a web-based application that allows users to input any message and instantly check whether it is Safe or Phishing.

The system analyzes key linguistic and structural features of the message and provides clear visual feedback:

  • 🟢 Green for Safe messages
  • 🔴 Red for Phishing messages

This helps users make quick and informed decisions before trusting suspicious messages.


🛠️ How we built it

Our project is a full-stack Machine Learning application:

  • Model: We trained a Random Forest Classifier using a dataset of thousands of labeled messages.
  • Feature Engineering: Extracted four key indicators:
    • Message length
    • Presence of URLs
    • Count of urgent keywords
    • Capitalization ratio
  • Backend: Built using Python and Flask to process user input and return predictions from a saved .pkl model.
  • Frontend: Developed with HTML and CSS, providing a clean and responsive user interface.
  • Deployment: Version-controlled with Git/GitHub and deployed live using Render.

🧠 Challenges we ran into

One major challenge was a feature mismatch error during deployment. While the trained model expected four numerical features, the live application initially sent raw text input.

We resolved this by restructuring our backend to perform real-time feature extraction before passing data to the model. Additionally, we learned the importance of repository hygiene by implementing a .gitignore file to keep the deployment clean and efficient.


🏆 Accomplishments that we're proud of

  • Successfully deployed a Machine Learning model as a live, cloud-hosted web application.
  • Built a complete end-to-end ML pipeline from training to deployment.
  • Accurately detected real-world phishing messages, including complex bank scam examples, during live testing.

📚 What we learned

Through this project, we gained hands-on experience with the *end-to-end Machine Learning life cycle *, including data preprocessing, feature engineering, model training, API integration, and cloud deployment.

We also strengthened our skills in version control, debugging production issues, and building user-focused security applications.


🚀 What's next for Phishing Detector

  • Improve accuracy using advanced NLP models such as TF-IDF or transformer-based approaches.
  • Support multiple languages to reach a wider user base.
  • Add browser extension and mobile app support for real-time protection.
  • Enhance the UI with detailed explanations on why a message is flagged as phishing.
Share this project:

Updates

Submission history