Inspiration

Wildlife sightings can tell us where an animal has been observed, but they don't necessarily tell us why an animal might prefer one location over another.

That led us to a question: What if we could use machine learning to understand the environmental conditions that make a location suitable for a specific species?

We wanted to build something that went beyond simply displaying wildlife observations on a map. Instead, we envisioned a tool that could evaluate actual environmental conditions and let users explore how roads, buildings, terrain, and other geographic features relate to wildlife habitat suitability.

That's how Wild-Locate started.

What it does

Wild-Locate is a machine-learning-powered habitat analysis platform that helps users investigate wildlife habitats using real environmental data.

Users select a species and choose a location on an interactive map. They can analyze either a single coordinate or an entire surrounding region with a radius of 10, 25, or 50 kilometers.

Wild-Locate generates a relative habitat suitability percentile, showing how the selected location compares to other locations evaluated by the model.

However, the main feature of Wild-Locate is Deep Dive.

Rather than stopping at a single suitability score, Deep Dive analyzes the surrounding region and identifies actual mapped roads, buildings, habitat areas, and terrain locations.

For each selected location, Wild-Locate evaluates environmental conditions and identifies which factors may support or limit habitat suitability. It also provides environmental comparisons, geographic information, and potential issues worth investigating further.

For example, someone studying bobcat habitats can investigate an individual road and examine nearby development, land cover, or terrain conditions that might be associated with lower habitat suitability.

Wild-Locate also supports training custom habitat models for additional mammal and reptile species.

Our predictions are relative measures of habitat suitability, not probabilities of animal presence. Deep Dive highlights environmental associations rather than proving that individual structures cause ecological harm.

How we built it

We developed Wild-Locate primarily using Python, combining machine learning with geospatial data processing and an interactive interface.

We integrated several environmental and wildlife datasets:

  • iNaturalist: Research-grade wildlife observations used to train species-specific models.
  • USGS National Land Cover Database (NLCD): Forest cover, developed land, wetlands, and impervious surfaces.
  • USGS 3D Elevation Program (3DEP): Elevation and terrain characteristics.
  • MassGIS / MassDOT: Road networks and hydrographic information in Massachusetts.
  • OpenStreetMap: Real geographic features, including individual roads, buildings, and habitat areas.

We first collect and filter wildlife observations, removing inaccurate, duplicate, or unsuitable records. We then generate background samples and extract environmental variables from the corresponding geographic locations.

Using these features, we train and evaluate Logistic Regression, Random Forest, and XGBoost models. We use spatial cross-validation to help assess performance across geographically separated locations.

For Deep Dive, we built an additional analysis pipeline that evaluates 81 geographic sample points within the selected region. It identifies neighborhoods worth investigating, retrieves real-world geographic features through OpenStreetMap, and evaluates the environmental conditions around individual locations.

To make the predictions more interpretable, we also compare environmental variables against reference conditions, measuring how individual changes affect the model's predicted suitability score.

We built the web interface using HTML, CSS, JavaScript, and Leaflet, with a Python backend handling the machine learning and geographic processing. We also developed a desktop interface using PyQt.

Challenges we ran into

One of our biggest challenges was integrating multiple geographic datasets.

Wildlife observations vary significantly in geographic coverage and accuracy. Some locations have hundreds of sightings, while others have almost none. We had to account for observation bias, filter unreliable records, and construct appropriate background samples.

Environmental data presented another challenge. Land cover, elevation, hydrography, and road datasets have different formats, resolutions, and coverage. Combining them into a consistent machine learning pipeline required extensive preprocessing and debugging.

Developing Deep Dive was particularly challenging. Retrieving actual road and building geometries, connecting them to environmental measurements, and displaying everything accurately on an interactive map introduced a new set of technical problems.

We also struggled with presenting our findings clearly. At first, the interface contained too many statistics and overlapping analyses. We eventually redesigned it to provide a simple initial assessment, with more detailed information available through Deep Dive.

Another important challenge was ensuring that our interpretations remained scientifically reasonable. We needed to distinguish between environmental associations and actual causal impacts.

Accomplishments that we're proud of

We're especially proud of transforming complex environmental datasets into an application that users can interact with directly.

Our biggest accomplishment is Deep Dive. Rather than only predicting habitat suitability across a region, Wild-Locate can connect model results to actual geographic features that users can locate and investigate.

We're also proud of implementing an end-to-end machine learning pipeline, from collecting wildlife observations and engineering environmental features to training models and generating interactive predictions.

Additionally, Wild-Locate supports multiple species, custom model training, and both point-based and regional analysis.

Seeing everything come together, especially being able to click on an actual road or building and investigate its surrounding environmental conditions, was one of the most rewarding parts of the project.

What we learned

Building Wild-Locate taught us about machine learning, geospatial analysis, environmental datasets, and full-stack development.

We gained hands-on experience with geographic APIs, spatial cross-validation, model interpretability, data preprocessing, and interactive mapping.

But our biggest takeaway was that building an accurate machine learning model is only part of the challenge.

Making the model's predictions understandable, reliable, and useful is just as important.

We learned to think critically about the limitations of our data and how to communicate uncertainty without making our application unnecessarily complicated.

What's next for Wild-Locate

We want to continue expanding Wild-Locate's geographic coverage and improving the depth of its habitat analysis.

Our next goals include incorporating more detailed environmental datasets, improving Deep Dive's ability to investigate individual geographic features, and supporting more species and regions.

We also hope to explore more rigorous methods for studying how potential changes to roads, development, and land cover could affect habitat suitability.

Ultimately, our goal is to make wildlife habitat analysis more accessible, interactive, and useful, helping people better understand the relationship between animals and their environments.

Built With

Share this project:

Updates

Submission history