Inspiration

As a botany student with research interests in stomatal biology and water-use efficiency, I kept coming back to the same practical gap: farmers usually only notice water stress once a plant visibly wilts by which point yield has already been lost. Soil-moisture sensors can catch it earlier, but they're expensive and impractical to install across an entire field. I wanted to ask a harder, more useful question: can water stress be detected remotely, using signals a drone or weather station can already capture, without needing soil access at all?

What it does

StomaSense predicts crop water-stress severity (on a 4-level scale) using only canopy structure (Leaf Area Index) and weather data (VPD, ET₀). It's built as a hybrid decision-support system: a Random Forest model does the prediction, and a Streamlit interface turns that prediction into a plain-language irrigation recommendation.

How we built it

The project went through several deliberate pivots:

  1. Dataset selection — I initially explored a synthetic plant-biosensor dataset, but its label was disease severity, not water stress. I switched to a real experimental dataset from a banana irrigation-deficit trial (Mendeley Data, Colombia), which has a native, ground-truthed water-stress label across four irrigation treatments.
  2. A red flag, caught early — My first model, trained with soil-moisture features included, scored 100% accuracy. Rather than celebrate, I treated this as a warning sign: in this experimental design, irrigation treatment directly determines soil moisture, so the model was essentially reading off the experimental setup, not learning biology.
  3. Reframing the problem — I deliberately excluded all soil-moisture-derived features and retrained using only canopy (LAI) and weather (VPD, ET₀) data the signals actually available without a soil sensor. This produced a more honest, harder, and more useful model.
  4. Building the interface — I built a Streamlit app around the retrained model, then deployed it on Streamlit Community Cloud so it's accessible without needing my local machine running.

Challenges we ran into

The near-perfect-accuracy trap: it would have been easy to ship the 100% accuracy soil-based model. Recognizing why that result was misleading rather than just accepting a good number was the most important technical decision in the project. Environment setup: as a biology researcher newer to full ML pipelines, getting Python, pip, and the local environment correctly configured (PATH issues, version matching between Colab and local) took real troubleshooting. Class imbalance: the four stress classes weren't evenly distributed, addressed using class_weight='balanced' in the Random Forest.

Accomplishments that we're proud of

  • A model that achieves 62% accuracy on a genuinely hard 4-class problem (vs. ~25% random baseline) using only remotely-observable signals.
  • A confusion matrix showing that nearly all errors occur between neighboring stress levels rather than extremes evidence the model learned a real physiological gradient, not noise.
  • A live, deployed, interactive prototype , not just a notebook.

What we learned

The most valuable lesson wasn't a specific tool or library it was learning to be suspicious of results that seem too good, and to trace why a model performs the way it does back to the actual experimental design and biology behind the data, rather than just optimizing a metric.

What's next for StomaSense

  • Extending training data to additional crop species beyond banana
  • Incorporating UAV-derived vegetation indices (NDVI, NDRE) alongside LAI for richer canopy signal
  • Field-testing the early warning thresholds against real-time deployment scenarios

Built With

Share this project:

Updates

Submission history