🛰️ ♳ Pelago

What they dump, we find.

Satellite-powered detection of water-adjacent dump sites, tracked over time.

License: MIT Python Frontend Build PRs Welcome Stars


The Pitch

Illegal and unregulated dump sites near rivers, lakes, and coastlines are some of the least-monitored pieces of the plastic-pollution problem — and some of the most damaging. Pelago finds them from space: a two-stage machine-learning pipeline screens Sentinel-2 satellite imagery at 10 m resolution, flags candidate sites that sit or drain toward water, runs every candidate through analyst validation, and monitors confirmed sites over time.

We turn open satellite data into an operational map of the places where plastic enters the world's waterways.


The Problem 🌏

Plastic doesn't start in the ocean. It starts on land, next to water.

Reality What it costs
🏞️ Rivers are highways for plastic A small number of river systems carry most ocean plastic; whatever is dumped near a river goes downstream.
🗑️ Dumpsites are unmonitored Most sites are informal, unregulated, or simply off any map — no government has a complete inventory.
🛰️ Satellites already see it Sentinel-2 revisits the entire land surface every few days at 10 m resolution. The imagery exists; it just isn't looked at systematically.
📋 Manual surveys don't scale Field audits are invaluable but cannot cover a watershed, let alone a continent, at the pace of the problem.

Pelago closes that gap — turning imagery that already exists into an inventory that governments, researchers, and cleanup organizations can act on.


The Solution 💡

Pelago finds water-adjacent dump sites with a two-stage ML pipeline. A pixel classifier scores every pixel for the spectral signature of waste; then a patch classifier, trained on weakly-labeled examples at scale, confirms whether the pixel detections form a real site. Only the intersection of both stages becomes a candidate — then analysts validate and confirm each one.

Stage What it does Why it matters
🧩 Pixel classifier Scores every Sentinel-2 pixel for waste spectral signature Casts a wide net over whole regions with spatial context removed
🧠 Patch classifier Confirms pixel detections against a weakly-trained site-level model Eliminates lookalikes (greenhouses, bare earth) that fool single pixels
✂️ Intersection filter Keeps only candidates that pass both pixel and patch agreement Dramatically cuts false positives before they cost analyst time
👁️ Manual validation Analysts review each candidate in high-resolution imagery Confirms / rejects / tiers every site with human-grade certainty
📈 Monitoring Confirmed sites get monthly contours, metadata, and history Tracks expansion and new activity over time

The result: a continuously updated, API-backed, map-first inventory of the world's water-adjacent waste sites.


Architecture ⚙️

┌──────────────────────────────────────────────────────────────────────────┐
│                         INGESTION                                        │
│  Sentinel-2 (10 m RGB + multispectral) → Descartes Labs catalog         │
│  Population-weighted tile generation → per-region download queue        │
└──────────────────────────────────┬───────────────────────────────────────┘
                                   │
┌──────────────────────────────────▼───────────────────────────────────────┐
│                         MODEL LAYER                                      │
│  Pixel classifier (spectral waste signature, per-pixel)                 │
│        │                                                                 │
│        ▼                                                                 │
│  Patch classifier (weakly-labeled ensemble over 28×28×24 windows)       │
│        │                                                                 │
│        ▼                                                                 │
│  Intersection filter (pixel ∩ patch agreement)                          │
└──────────────────────────────────┬───────────────────────────────────────┘
                                   │
┌──────────────────────────────────▼───────────────────────────────────────┐
│                         SITE DETECTION                                   │
│  Candidate generation (blob detection on scored tiles)                  │
│        │                                                                 │
│        ▼                                                                 │
│  Manual validation (analyst review of high-res imagery)                 │
│        │                                                                 │
│        ▼                                                                 │
│  Confirmed sites (confirmed / industrial / uncertain / negative)        │
└──────────────────────────────────┬───────────────────────────────────────┘
                                   │
┌──────────────────────────────────▼───────────────────────────────────────┐
│                         METADATA + MONITORING                            │
│  Contour generation (per-site footprint over time)                      │
│  Metadata enrichment (centroids, addresses via Nominatim)               │
│  Results pushed to the Pelago API (Go loader)                           │
└──────────────────────────────────┬───────────────────────────────────────┘
                                   │
┌──────────────────────────────────▼───────────────────────────────────────┐
│                         PRESENTATION                                     │
│  Web map (Next.js + MapLibre)  •  Alerts  •  Monitoring dashboard       │
│  Open GeoJSON/CSV export  •  API access                                 │
└──────────────────────────────────────────────────────────────────────────┘
  1. Ingestion. The pipeline pulls Sentinel-2 scenes through Descartes Labs, constrained to a population-weighted tile grid so compute goes where people — and waste — actually are.
  2. Model layer. A pixel classifier scores the spectral signature of waste across each tile. A patch classifier, trained with weakly-labeled data and ensembled across seeds, decides whether a scored region is genuinely a site. Candidates require agreement between both stages.
  3. Site detection. Scored regions become candidate sites through blob detection, then pass through analyst manual validation before being labeled confirmed.
  4. Metadata + monitoring. Confirmed sites receive contours (footprint and area over time) plus enriched metadata — centroid, Nominatim address, catchment context.
  5. Presentation. The web application renders sites on an interactive MapLibre map with confidence layers, monitoring timelines, and alert feeds — with the inventory exposed for export and programmatic access.

Key Features ✨

  • 🌍 Global satellite coverage — powered by open Sentinel-2 data with a population-weighted detection footprint.
  • 🧠 Two-stage ML pipeline — pixel classifier → patch classifier → intersection filter for high-precision detection.
  • 🎓 Weakly-supervised training — patch labels mined at scale from spectral + spatial priors rather than hand-annotated imagery.
  • 🎯 Confidence-tiered detections — confirmed, industrial, uncertain, and negative categories keep the inventory honest.
  • 🔭 Continuous monitoring — confirmed sites are revisited and contoured over time (rolling median masks over a monthly time series).
  • 🗂️ Open data export — candidate sites, validated points, and metadata as standard GeoJSON/CSV.
  • 🔌 API access — a Go loader service pushes model outputs into the Pelago API for the web app and integrations.
  • 🛰️ Sentinel-2 native — detections grounded in real 10 m multispectral observations, not modeled estimates.

Tech Stack 🧰

Layer Technology Purpose
🛰️ Satellite data Copernicus Sentinel-2 10 m multi-spectral imagery of the full land mass
☁️ Cloud imagery access Descartes Labs Bulk scene search, download, deployable inference
🧠 Modeling Python · TensorFlow · scikit-learn Pixel classifier, patch classifier, feature pipelines
📐 Geospatial GeoPandas · Rasterio · Shapely · PyProj Vector and contour processing
⚙️ MLOps / Deploy Descartes Deploy endpoints Distributed candidate detection and contour runs
🔌 Backend API Go · Pelago loader (loader/) Ingest model outputs, serve site data
🌐 Frontend Next.js · MapLibre · TypeScript Interactive globe, site detail, monitoring dashboard
💾 Data formats GeoJSON · CSV · HDF5 Interchange between models, validation, and API

Pipeline in Detail 🔬

Model Training

Module Inputs Role
create_pixel_dataset · create_spectrogram_dataset Raw Sentinel-2 tiles (+ 3-month minimum composites, 6-month offsets) Assemble per-pixel training chips and temporal spectrograms
train_pixel_classifier · train_spectrogram_classifier Pixel-level crops Learn the spectral waste signature; emit a temporal pixel model
train_patch_classifier Weak labels + pixel scores Learn site-level structure; ensemble, SVM, and 1-px variants trained at scale

Notebooks: create_pixel_dataset.ipynb, create_spectrogram_dataset.ipynb, train_pixel_classifier.ipynb, train_spectrogram_classifier.ipynb, Train Patch Classifier (Weak Labeling — Ensemble, SVM, 1px, LARGE).ipynb. Trained artifacts live under models/ (e.g., spectrogram_v0.0.11, weak_labels_28x28x24).

Site Detection

Module Role
generate_populated_dltiles Compute a population-weighted tile grid for a region
descartes_spectrogram_run_withpop Deploy pixel + patch inference over the region on Descartes
descartes_candidate_detect Run blob detection (scikit-image DoH) to extract candidate sites from scored tiles
query_patch_classifier Keep only candidates that pass the patch-classifier intersection
validate_candidate_sites Analyst review of candidates → confirmed / industrial / uncertain / negative

Metadata & Monitoring

Module Role
generate_metadata Enrich confirmed sites with centroids, addresses, and catchment context
descartes_contour_run Generate per-site contours on Descartes (boundary and area over time)
loader Push enriched model outputs into the Pelago API for the web app

Performance & Metrics 📊

Metric Value
🌎 Confirmed positive sites 4,700+ validated across operating regions
🏷️ Labeled validation dataset 19,900+ features (positive / negative / industrial / uncertain)
🗺️ Largest tracked sites Top 100 largest sites by area, maintained in-site inventory
📦 Regional exports 38+ interactive map exports under docs/
📏 Spatial resolution 10 m (Sentinel-2)
🌍 Detection footprint Population-weighted global tile grid
🔎 Validation workflow Analyst-reviewed, confidence-tiered
🔄 Refresh cadence Recurring on each Sentinel-2 revisit cycle
📈 Detection recall (Indonesia) 80% in high-sensitivity mode; 40% in low/medium
✅ Candidate precision (SE Asia) 53% of pixel∩patch candidates confirmed as waste by analysts

The system found 996 confirmed waste sites across Southeast Asia — nearly three times the number recorded on OpenStreetMap.


How We Built It 🛠️

The hard part was never the satellite. The hard part is that the spectral signal of waste is subtle, varied, and noisy — and hand-labeled examples are scarce. A single CNN trained on a handful of known sites overfits and under-generalizes. So we designed the system around the data constraints.

First, we went per-pixel. Operating on individual pixels instead of spatial patches multiplies the usable training signal from each known site and suppresses spatial overfitting. A temporal component — 3-month minimum composites offset by 6 months — suppresses spectrally-similar backgrounds like tilled fields and seasonal vegetation. Positive-class pixels are filtered through an NDVI threshold at 0.4 to keep vegetated pixels out of the waste class.

Then we added the second network. Because pixels alone can't distinguish a dump site from a greenhouse roof, a patch classifier — trained on weakly-labeled examples at scale and ensembled across seeds — validates every pixel-classifier candidate. Training data is continuously augmented with prior true and false positives, including via semi-supervised distillation.

Inference runs as deployable Descartes Labs jobs: the platform breaks each region into sub-tiles processed in parallel, with the patch classifier convolved at an 8-pixel stride to catch sites that straddle window boundaries. Candidate blobs are detected with scikit-image's Deterministic-of-Hessian detector, with three sensitivity regimes. And the validation workflow — analysts reviewing candidates against very-high-resolution imagery alongside Google Street View and Planet data — is what keeps false positives out of the confirmed inventory.


Challenges We Ran Into 🧗

Challenge How we solved it
🎲 Spectral waste signal is subtle and noisy Pixel-level classification amplifies signal per site; temporal spectrograms suppress seasonal lookalikes
🏷️ Patch labels are scarce Weakly-supervised labels mined from spectral + spatial priors; semi-supervised distillation on unlabeled data
🌫️ Cloud and haze fooled early models Minimum-composite (instead of median) over 3-month windows removes bright haze and wispy clouds
🖼️ Lookalikes (greenhouses, bare earth) pass pixel tests Patch classifier acts as a spatial validator; only pixel ∩ patch agreement becomes a candidate
🧮 Distributed contour generation Rolling median masks (8-frame) filter outlier predictions before findContours on Descartes
🔄 Keeping the API in sync with model outputs Go loader computes centroids, geocodes via Nominatim, and reconciles newly-generated sites against existing API records
🎯 Candidate tuning without ground truth Three sensitivity modes (threshold × min_sigma) tuned on surfaced candidates rather than an assumed optimum

Accomplishments We're Proud Of 🏆

  • 🌏 Continental-scale pipeline — deployed across South & Southeast Asia, the Mediterranean, Africa, and South America from a single codebase.
  • ✂️ Two-stage intersection that cuts false positives — candidate volume is reduced by the pixel ∩ patch filter before any analyst time is spent.
  • 🗂️ Open, reproducible data — every candidate, validated point, contour, and metadata table exported as standard GeoJSON/CSV.
  • 📝 Published methodology — "Satellite monitoring of terrestrial plastic waste" manuscript with quantitative validation of both classifiers.
  • 🔎 Validated at scale — 4,700+ confirmed sites, a 19,900+ feature labeled dataset, and top-100 site inventory maintained in-repo.
  • 🤝 A real team effort — built as a single integrated pipeline by Flowthread for NextStep Hacks 2026.

What We Learned 📚

  • 🛰️ Satellite ML — spectral signatures of waste are real but subtle; temporal compositing is what makes them usable.
  • 🎓 Weakly-supervised training — mined labels beat scarce hand labels when you architect the model around them.
  • 🗺️ Geospatial pipelines — GeoJSON/CSV as the interchange format makes every stage inspectable and every export open.
  • ⚙️ Distributed inference on Descartes — population-weighted tiling and deployable jobs make continent-scale runs tractable.
  • 📏 Working with Sentinel-2 at 10 m — single pixels lie; time-averaging, masking, and ensemble agreement is what earns trust.

What's Next 🚀

  • [x] Two-stage pixel → patch detection pipeline
  • [x] Operational candidate generation on Descartes Labs
  • [x] Analyst validation workflow and confidence-tiered inventory
  • [x] Contour generation and site-level monitoring
  • [x] API ingestion loader and open GeoJSON/CSV export
  • [x] Region expansions: South & Southeast Asia, Mediterranean, Africa, South America
  • [ ] Expanding detection to new river basins and coastal watersheds
  • [ ] Adding higher-cadence revisit monitoring for high-risk sites
  • [ ] Integrating additional Sentinel-2 band products for finer material discrimination
  • [ ] Publishing a public API endpoint for alert subscriptions

Links 🔗


What they dump, we find.

Built for NextStep Hacks 2026 — Earth Forward 🌱

Built With

Share this project:

Updates

Submission history