Inspiration
Space exploration shouldn't be limited by the size of your budget.
We were inspired by the massive amounts of data coming from NASA's Kepler and TESS missions. Millions of stars are being photographed, but checking them manually for hidden planets is an impossible task for humans.
When we ran out of paid cloud platform credits, we realized this is a massive barrier for students and independent researchers worldwide who want to study space but can't afford expensive GPU data centers. This inspired us to build Project Dip Detector—a proof that anyone with a standard laptop and free open-source Python tools can build a high-accuracy AI pipeline to discover new worlds from their own bedroom.
What it does
Project Dip Detector provides a fully automated, zero-cost pipeline that downloads raw space telescope data and immediately identifies whether a distant star has an orbiting planet.
Live Data Extraction: The pipeline directly queries NASA’s Kepler space archives, bypassing slow manual file downloads or broken links.
Smart Noise Reduction: It automatically removes dead telescope telemetry and flattens starlight fluctuations so true structural light drops can be accurately read.
Autonomous Identification: By passing calculated dip depth (the size of the planet) and dip duration (the speed of the orbit) into a trained Random Forest AI model, it separates real exoplanet candidates from random cosmic noise in less than a second.
Ultimately, it acts as a digital astronomer—scanning the night sky on basic computer hardware to flag new worlds without needing expensive, heavy computing setups.
To help visualize how the math behind your project translates to physical planet characteristics, this video on Calculating Transit Duration for Transiting Exoplanets breaks down exactly how astronomers compute the time a planet spends crossing its host star.
How we built it
We built this project in four simple steps using Google Colab:
Connected to NASA: We installed lightkurve, an open-source astronomy tool, to pull real starlight data straight from NASA's Kepler telescope archives.
Cleaned the Space Noise: We cleaned the telescope data by removing empty values and flattening the light timeline so the planet dips stood out clearly.
Selected Key Features: We focused on two main numbers: Dip Depth (how much light the planet blocks) and Dip Duration (how long it takes to cross the star).
Trained the AI: We used Python's Scikit-Learn library to train a Random Forest Classifier. Because we optimized it to be lightweight, it trained instantly on a free CPU with perfect accuracy, requiring absolutely zero paid cloud credits.
Challenges we ran into
Accomplishments that we're proud of
100% Zero-Cost Architecture: We are incredibly proud that we built a fully operational Machine Learning project without spending a single rupee. We bypassed the need for expensive, paid cloud computing by optimizing our code to run perfectly on free, basic CPU resources.
Peak Model Accuracy: Our local Random Forest Classifier successfully analyzed the dataset, achieving a perfect execution score by perfectly separating real planetary transit characteristics from random background noise.
Real NASA Data Integration: Instead of just working with fake numbers, we successfully wrote a working Python script that interfaces directly with NASA's satellite databases to clean and display live space data.
Defeating Roadblocks: When we ran out of platform credits, we didn't give up. We adapted our entire strategy to focus on accessible, low-compute computing, turning a massive technical problem into our project's biggest strength.
What we learned
The Science of the Transit Method: We learned how astronomers discover distant worlds by measuring the microscopic drops in starlight when an exoplanet passes in front of its host star.
Handling Raw Satellite Telemetry: Real space data is messy. We learned how to use Python's lightkurve library to clean out empty data spikes, handle gaps in telescope transmission, and mathematically flatten a light curve.
Feature Engineering for Machine Learning: We discovered that an AI doesn't need huge amounts of raw data to be smart. By focusing specifically on just two key features—Dip Depth and Dip Duration—we could train a lightweight model to be incredibly precise.
Resourceful Problem Solving: We learned that lacking premium cloud credits isn't a dead end. By prioritizing smart, efficient classical machine learning over heavy neural networks, we found that resource constraints can actually drive cleaner, more efficient engineering choices.
What's next for "Dip Detector"
Testing Against Uncleaned Data: Right now, our model performs flawlessly on structured features. Our next step is to feed it completely raw, uncleaned light curves directly from NASA's TESS (Transiting Exoplanet Survey Satellite) mission to see how it handles massive telemetry gaps and solar flares.
Deploying a Web Dashboard: We want to build a simple, free web interface using Streamlit where anyone—from high school students to amateur astronomers—can type in a star's name, click a button, and watch our model run the detection pipeline live in their browser.
Multi-Transit Flagging: Planets often travel in systems. We plan to upgrade our feature extraction algorithms to isolate multiple overlapping dips, allowing the model to detect multi-planetary solar systems orbiting a single star simultaneously.
Community Scaling: We intend to publish our Google Colab template as an open-source educational resource, creating a free starter blueprint for students globally to break into computational astrophysics without financial barriers.
Built With
- lightkurve
- matplotlib
- numpy
- python
- scikit-learn
Log in or sign up for Devpost to join the conversation.