Inspiration

Astronomers don't see exoplanets directly — they infer them from a faint, periodic dip in starlight as a planet crosses its star. We wanted to tackle the exact classification problem NASA scientists face when vetting Kepler candidates, using real observational data instead of a synthetic dataset — and to do it honestly, without letting the model quietly "cheat" off NASA's own vetting labels.

What it does

Classifies Kepler Objects of Interest (KOIs) as CONFIRMED, CANDIDATE, or FALSE POSITIVE using real transit and stellar measurements from the NASA Exoplanet Archive's Kepler cumulative table (9,564 rows, 140 columns).

Model Accuracy Macro F1 Weighted F1 Macro Recall
Random Forest, raw-signal (primary) 0.855 0.823 0.853 0.820
Random Forest, + NASA vetting flags (comparison only) — 0.896 0.920 0.890

The gap between these two rows is the point — see "Why two models" below.

Why two models

Several columns in the Kepler catalog are NASA's own near-final verdicts, not raw measurements — koi_pdisposition, and identifier columns like kepler_name (named planets are almost always already confirmed). Including these would let a model quietly "copy the answer" instead of learning real signal. We excluded them entirely from the primary model, and isolated NASA's diagnostic false-positive flags (koi_fpflag_*) into a separate, clearly labeled comparison model rather than folding them in silently. The primary 0.855 accuracy / 0.823 macro F1 result is therefore a genuine estimate of what's learnable from raw transit and stellar signal alone — not a NASA verdict in disguise.

How we built it

Pipeline:

Raw KOI CSV (9,564 rows × 140 cols)
        │
        ▼
Leakage Audit ──► koi_pdisposition, koi_score, kepler_name, IDs → DROPPED
        │          koi_fpflag_* (vetting flags) → isolated to comparison model
        ▼
EDA ──► class imbalance, missingness, correlation, outlier scan
        │
        ▼
Feature Engineering ──► transit SNR proxy, radius-ratio consistency,
        │                log-transforms, error-to-value ratios
        ▼
Stratified 5-Fold CV ──► imputer fit fresh INSIDE each fold (no leakage)
        │
        ▼
Model Training ──► Random Forest vs. XGBoost, tuned
        │
        ▼
Evaluation ──► Confusion Matrix · PR/ROC curves · CV stability
        │
        ▼
Explainability ──► SHAP values · plain-language summary
Stage Key decision
Leakage audit Dropped columns that are NASA's own near-final verdicts, not raw measurements
EDA Quantified 50.6% / 28.7% / 20.7% class split before choosing a resampling strategy
Feature engineering Built transit-physics-aware features, not generic polynomial terms
Validation Stratified 5-fold CV, imputer refit per fold to prevent validation leakage
Explainability SHAP + a plain-language paragraph, not just a feature-importance plot

Stack: Python, scikit-learn, XGBoost, SHAP, pandas, NumPy, Jupyter, Matplotlib, Seaborn

Challenges we ran into

The biggest challenge wasn't modeling — it was realizing several catalog columns are NASA's own near-final verdicts rather than raw measurements. kepler_name is especially subtle: named KOIs are almost always already confirmed, so it's an almost-perfect leak that doesn't look like one at first glance. Designing a primary/comparison model split, instead of just picking whichever number looked better, was the harder and more important decision.

Accomplishments that we're proud of

Accomplishment Why it matters
Zero-error, top-to-bottom reproducibility Restart Kernel & Run All succeeds with no manual fixes
Fold-safe imputation Missing-value stats computed only from training folds, never validation folds
Two-model leakage comparison Shows exactly how much apparent accuracy inflation comes from NASA's own flags
Physically-grounded features Every engineered feature ties to real transit photometry, not blind feature crossing

What we learned

The most important raw-signal features were koi_dikco_msky, koi_dicco_msky, koi_depth_rel_err1, koi_model_snr, koi_ror, and koi_depth_mean_rel_err — these make physical sense: false positives often show unusually large inferred radii, high impact parameters, inconsistent radius-depth relationships, or extreme signal properties. In plain terms: the model checks whether a dip in starlight is the right size, lasts a believable amount of time, repeats in a stable orbit, and matches what we already know about the star. When it doesn't line up, that's a red flag for a false positive rather than a real planet.

What's next for Leakage-Free Kepler Exoplanet Classifier

Extending the same leakage-audit methodology to TESS Objects of Interest (TOI) data, and testing whether the transit-consistency features generalize across missions rather than being Kepler-specific artifacts.

Built With

Share this project:

Updates

Submission history