Inspiration

I wanted to learn how machine learning can be used to solve real scientific problems. The idea of helping identify planets outside our solar system using NASA's Kepler data really interested me, so I decided to build this project.

What it does

My project classifies Kepler exoplanet candidates into three categories: CONFIRMED, CANDIDATE, and FALSE POSITIVE. I compared multiple machine learning models, selected the best one, and used SHAP to understand why the model made its predictions.

How I built it

I started by cleaning and exploring the dataset. After handling missing values and preparing the data, I trained and compared Random Forest, XGBoost, CatBoost, and LightGBM. LightGBM gave the best performance, so I selected it as my final model. I also used SHAP to explain the model's decisions.

Challenges I ran into

The biggest challenge was that I had very little experience with machine learning before starting this project. I spent time understanding the dataset, choosing the right features, comparing models, and making sure the model was not relying on misleading information. Learning SHAP and explaining the model's predictions was also challenging but rewarding.

What I learned

This project taught me how to build a complete machine learning pipeline, compare different models, evaluate their performance, and explain predictions using SHAP. It also helped me understand how machine learning can support real scientific research.

Built With

Share this project:

Updates

Submission history