Inspiration
Crop yield is one of the most consequential numbers in the U.S. agricultural economy, yet weather-driven variability makes it notoriously hard to forecast early in a growing season. We wanted to see how far a weather-only signal could actually take us, and whether a modern lakehouse stack (Databricks, Unity Catalog, Genie) could turn a static prediction model into something a farmer, analyst, or policymaker could actually converse with. The Databricks x UW Hackathon gave us the perfect forcing function to build that idea end to end in a short window.
What it does
TerraCast predicts county- or state-level corn and soybean yield (bushels per acre) from growing-season weather features, using six regional LightGBM models split by crop and geography (corn and soybeans, each across midwest, plains, and south regions). On top of the prediction engine sits a Databricks Genie AI assistant that lets users ask natural-language questions over the merged weather and yield dataset and over the prediction history, translating them into SQL under the hood. Every prediction is logged to a Delta audit table, and a lightweight web UI lets users pull historical features for a given county, year, and crop to prefill their inputs.
How we built it
We merged NOAA growing-season weather data with USDA county-level corn and soybean yield records spanning 2010 to 2024, then engineered 45 features capturing the weather signal relevant to each crop. Rather than training one global model, we trained six separate LightGBM boosters, one per crop-region combination, since weather sensitivity and baseline yields differ meaningfully across the Midwest, the Plains, and the South. The whole pipeline runs on Databricks and Spark, with the merged dataset and a predictions log stored as Unity Catalog Delta tables. On the serving side, a FastAPI app exposes prediction, feature-lookup, and chat endpoints, deployed as a Databricks App; it deliberately never spins up a Spark job on the request path, instead reading and writing through the Databricks SQL Statement API for low-latency responses. The Genie Space is configured with dataset-specific instructions so it can answer natural-language questions about the underlying data and about past predictions.
Challenges we ran into
The biggest challenge was setting realistic expectations for a weather-only model. Weather explains a meaningful share of year-to-year yield variation, but soil quality, genetics, and farming practices matter too, so we had to be careful about what \R^2\ was actually achievable and honest about the ceiling of a weather-only approach. Getting the validation methodology right was just as important: naive random train/test splits leak information across years, so we moved to a leave-one-year-out scheme to get a trustworthy read on out-of-sample performance. On the engineering side, coordinating large-scale collaborative feature engineering under hackathon time pressure, and then integrating six independently trained regional models into a single coherent serving and logging pipeline, took real effort to get right without slowing down the app.
What we learned
Splitting a single problem into regionally specialized models can outperform one global model when the underlying weather-yield relationship genuinely differs by geography, but it also multiplies the operational surface area you have to manage in production. We also came away with a much sharper sense of how validation design, not just model choice, determines whether a yield-prediction result is actually trustworthy.
Built With
- databricks
Log in or sign up for Devpost to join the conversation.