Inspiration

Wild-Locate started with a question we kept coming back to: what actually makes a place suitable for wildlife?

A wildlife sighting gives you a point on a map. It does not tell you much about the area around that point: how much of it is forested, whether there is water nearby, how developed it is, what the terrain looks like, or how close it is to roads.

We wanted to connect those two pieces. Wild-Locate combines public wildlife observations with environmental data so that someone can pick a location and explore why a model considers it more or less suitable for a species.

What It Does

Wild-Locate is a desktop application that estimates relative habitat suitability for a mammal at a selected location.

A user can click anywhere on the map or enter coordinates, choose a species, and run an assessment. The result includes:

  • A suitability score
  • Its percentile relative to comparison locations
  • A suitability category
  • The environmental features at that point
  • Information about the model
  • The number of wildlife observations used to train it

Assessments can also be exported as JSON.

The Massachusetts version includes models for:

  • Bobcat
  • Coyote
  • Fisher
  • Red fox
  • North American river otter

Users are not limited to those species, though. They can train additional mammal models inside the application, inspect the validation results, and decide whether to enable a model for future assessments.

We also added experimental support for Florida and Arizona. Those states use their own regional datasets and model configurations rather than treating a model trained for one region as interchangeable with another.

One distinction is especially important when interpreting the results. A suitability percentile is not the probability that someone will see an animal there. It only describes how that location compares with the locations used by the model.

Wild-Locate also has a Habitat Insights tool for experimenting with individual features. For example, a user can see how the model output changes when forest cover is increased or impervious surface is reduced. These are controlled experiments on the model, not claims that making that environmental change would cause the same effect in the real ecosystem.

How We Built It

We built Wild-Locate in Python. The desktop interface uses PyQt6, while the interactive map is built with Leaflet and runs inside Qt WebEngine.

The actual analysis code is kept separate from the interface. The desktop application, command-line tools, and FastAPI endpoints all use the same underlying analysis layer, so an assessment should behave the same way regardless of how it is called.

For observations, we use Research Grade iNaturalist records. Those observations are combined with environmental data from sources including:

  • NLCD land cover
  • USGS 3DEP elevation
  • Massachusetts road datasets
  • Massachusetts hydrography datasets

GeoPandas, Rasterio, Shapely, and PyProj handle most of the geographic processing, coordinate transformations, and feature extraction.

For every training location, the pipeline measures environmental conditions in the surrounding area and turns them into model features.

The training system compares logistic regression and random forest models, with optional XGBoost support. Evaluation uses five-fold spatial cross-validation rather than an ordinary random split. Nearby locations stay in the same geographic block, which makes it harder for a model to receive an unrealistically high validation score simply because very similar locations appeared in both its training and validation data.

Models are selected primarily using average precision.

Training can also take long enough that running it directly through the interface would be a problem, so it happens in a separate process. The GUI receives progress updates and the user can cancel a run without freezing the rest of the application. Downloaded environmental data is cached for reuse, while models created by users remain associated with their local accounts.

Challenges We Ran Into

The geospatial data was probably the messiest part of the project.

Different datasets came with different coordinate systems, resolutions, file formats, coverage areas, and definitions of missing data. They could not simply be downloaded and fed into the same function. We had to make the pipeline transform them into a consistent form and also check downloaded raster tiles before using them.

Adding Florida and Arizona exposed another problem: the assumptions that worked for Massachusetts could not just be carried over to another state. We introduced separate regional feature schemas and validation checks so that, for example, a model trained using Massachusetts features cannot silently be used with incompatible Arizona data.

The wildlife observations introduced a different kind of problem. iNaturalist tells us where an animal was observed; it does not provide a matching dataset of places where that animal definitely was not present. Our negative examples therefore have to be sampled as background locations rather than treated as confirmed absences. We accounted for that distinction during training and also added checks for spatial folds that could not produce a valid training or validation set.

Then there was the application itself. Downloading and processing geographic data, extracting features, and training several candidate models can be expensive. We had to keep all of that work from locking up the interface while still giving the user useful progress information and a way to cancel it.

Finally, we had to decide how much interpretation the application should provide. A model can easily output a number to several decimal places; making sure that number is not presented with more certainty than it deserves was harder.

Accomplishments That We're Proud Of

  • Wild-Locate connects the full workflow in one application: wildlife observations, environmental data processing, model training, validation, mapping, and individual habitat assessments.
  • Users can train their own mammal models without having to build the geospatial and machine-learning pipeline themselves.
  • The application includes five Massachusetts species out of the box, along with experimental regional support for Florida and Arizona.
  • Every result comes with context, including its comparison percentile, environmental measurements, model information, and training-observation count.
  • Habitat Insights makes it possible to inspect how the model responds to individual environmental changes instead of treating its score as a black box.
  • Most analysis stays local, while environmental downloads are cached so the same large datasets do not need to be fetched repeatedly.

What We Learned

The biggest lesson was that choosing the machine-learning algorithm was only one part of habitat modeling.

How observations are collected matters. How background points are sampled matters. Geographic leakage between training and validation data matters. A model with a higher score under a bad evaluation setup can be less trustworthy than a simpler model evaluated properly.

We also spent more time than expected thinking about what the output should actually mean to a user. Showing a number is easy. Explaining what population that number is being compared against, what data produced it, and what conclusions should not be drawn from it is much more important.

That changed how we built Wild-Locate. Model information, validation results, environmental measurements, and observation counts are part of the result instead of being hidden behind the final score.

What's Next for Wild-Locate

The next step is external validation. We want to compare Wild-Locate's predictions against independent ecological datasets rather than relying only on held-out iNaturalist observations.

We also want to investigate observation bias more closely, especially in areas where iNaturalist activity is concentrated around roads, trails, or population centers. The Florida and Arizona models need more evaluation before we consider their support comparable to Massachusetts.

On the application side, we want to make environmental downloads less cumbersome, make location-to-location comparisons easier, and add more species and regions when there is enough data to support them responsibly.

Built With

Share this project:

Updates

Submission history