Inspiration
Singapore's dengue control is among the best in the region, yet outbreaks keep swinging hard. In 2020 there were 35,315 cases and 32 deaths, the worst year on record. Cases then fell to 5,258 in 2021 and jumped to 32,173 in 2022, a sixfold rise in one year. NEA already runs hundreds of thousands of inspections a year (about 841,000 from January to November 2022), so we realised the real gap is not effort, it is timing.
A cluster polygon shows where dengue already was. Published Singapore research suggests weather carries an earlier signal: Hii et al. (2012) built a model on temperature and rainfall alone that gave 16 weeks of warning with under 3% false alarm for the 2011 outbreak. Those models work largely at national level, but inspection teams work by neighbourhood. We wanted to bring that early warning down to the planning area.
What it does
DengueRadar turns live weather and cluster data into a ranked, explained dengue risk rating for every planning area, two weeks ahead.
- Map: toggle between current cluster intensity and the two week forecast
- Ratings: each area is Low, Medium or High relative to its own history, so dense towns are not permanently red
- Reasons: click an area to see plain language drivers, such as heavy rain three weeks ago, a long warm spell and a growing cluster next door
- Scorecard: an in app comparison of our model against a recent history baseline
- Genie: planners can ask questions like which areas rose most this week
How we built it
Everything runs on Databricks, from raw feeds to the demo.
- Ingestion: Auto Loader and Lakeflow bring the 5 minute rainfall and temperature readings and daily cluster snapshots into bronze Delta tables
- Silver: data quality expectations catch bad sensor readings, and station readings are mapped to planning areas with inverse distance weighting
- Gold: one row per planning area per week, with weekly rainfall totals, rainy days, longest warm run, neighbouring cluster burden, season flags and weather lags up to 8 weeks
- Model: gradient boosted trees predict cases two weeks ahead, benchmarked against a Poisson regression and a last 4 week average. Every run is tracked in MLflow, including a feature set ablation (history only, then plus weather, then plus neighbours)
- Serving: a daily scoring job writes the forecast table that the Databricks App map, AI/BI dashboard and Genie read
- Governance: Unity Catalog gives lineage from raw feed to each rating
Formally, for planning area \(a\) in week \(t\), the model learns
$$ \hat y_{a,t+2} = f(\mathbf x_{a,t}) $$
where \(\mathbf x_{a,t}\) includes rainfall \(R_{a,t,k}\) and temperature \(T_{a,t,k}\) at \(k\) weeks of lag. The forecast is converted to a rating using each area's own history:
$$ r_{a,t} = \Pr(Y_a \le \hat y_{a,t+2}) $$
with Low if \(r_{a,t} < \tau_1\), Medium if \(\tau_1 \le r_{a,t} < \tau_2\) and High if \(r_{a,t} \ge \tau_2\). Explanations come from SHAP values, which split each prediction into additive contributions:
$$ \hat y = \phi_0 + \sum_j \phi_j $$
We validated with walk forward splits in time, never random splits, to avoid leaking the future.
Challenges we ran into
- The API did not match its spec. Our first ingestion run crashed because station records lacked the field the OpenAPI document promised, so we made the parser defensive and inspected raw responses first.
- No public planning area history. The cluster feed is a snapshot of active clusters, not a time series, so we started snapshotting it daily into Delta to build our own history.
- Mismatched cadences. Weather arrives every 5 minutes, but cases and clusters arrive daily or weekly, so we are honest that the forecast refreshes daily, not every 5 minutes.
- A blind spot we cannot fix with weather. Serotype shifts matter: serotype 3 rose from 19.1% to 40% of cases between January and March 2026. We use recent case history to absorb some of this and name the limitation openly.
Accomplishments that we're proud of
- A governed pipeline from live feeds to a ranked weekly action list, with lineage we can show end to end
- A model that beats the recent history baseline: MAE [X]% lower, recall on the High class of [Y]%
- A targeting result that is easy to remember: if teams inspect only the top 5 ranked areas, they cover [Z]% of the next two weeks' cases versus [W]% under the baseline ranking, where
$$\text{Coverage}@K = \frac{\sum_{a \in \text{Top}K} y_{a,t+2}}{\sum_{a} y_{a,t+2}}$$
- Plain language reasons behind every rating, so a planner can check the map instead of trusting it blindly
What we learned
- Weather signal arrives with lags and is patchy across the island, so mapping stations to areas mattered more than we expected.
- Evaluation design decides credibility: time based splits, real baselines and a ranking metric planners care about.
- Admitting what the model cannot see, such as serotype shifts, made the project more trustworthy, not less.
- Building the data foundation took most of the time, and snapshotting live feeds from day one pays off.
What's next for DengueRadar
- Threshold alerts when an area's risk crosses into High, using Databricks SQL alerts
- Environmental layers such as hawker centre density and green space
- Humidity as a feature, since absolute humidity has been a strong predictor in Singapore modelling
- Serotype data, if it becomes available, to cover our biggest blind spot
- A pilot with a town council or outreach team to test whether ranked lists change where inspections go
Built With
- ai/bi-dashboards
- apache-spark
- auto-loader
- data.gov.sg
- databricks
- databricks-apps
- delta-lake
- genie
- geojson
- geopandas
- lakeflow
- lightgbm
- mlflow
- nea-open-data
- pandas
- pyspark
- python
- rest-api
- shap
- sql
- unity-catalog

Log in or sign up for Devpost to join the conversation.