DengueRadar — Our Project Story
What inspired us
Singapore is one of the world's most liveable cities, yet every few years dengue returns with devastating force — 35,261 cases in 2020, peaking at 1,791 in a single week, and 32,130 again in 2022. The numbers are not abstract to us: they are neighbours, classmates, and grandparents in our own planning areas.
What struck us most was when people learn they are at risk. The National Environment Agency's (NEA) public cluster map and myENV alerts tell you where dengue is today — after people are already sick and Aedes mosquitoes have been breeding for weeks. By that point, prevention is already late. The Aedes life cycle can be as short as seven days, so a two-week warning covers two full breeding cycles — enough time for a resident to empty a flowerpot plate, clear a gutter, or check on an older neighbour.
That gap became our problem statement: NEA's public dashboard shows today's clusters; nothing public shows where risk is heading over the next two weeks, why, and what to do about it. We built DengueRadar to close that gap — to move dengue prevention from reactive to predictive.
How we built it
DengueRadar is a near-real-time forecasting pipeline on Databricks, forecasting each of Singapore's 55 planning areas two weeks ahead.
The pipeline. Data flows from data.gov.sg and NEA APIs through Lakeflow Jobs (a daily cluster archive and incremental weather loading) into Delta Lake tables organised bronze → silver → gold, governed in Unity Catalog with lineage. A LightGBM model — tracked and registered in MLflow — forecasts each area's cluster risk, with SHAP surfacing the top three drivers behind every forecast. The result lands in an AI/BI map and a Databricks App, with Genie answering plain-English questions on the same tables.
The model. We frame the problem as both regression and classification. Our "Rising" signal flags an area when the next fortnight's cluster cases clear a threshold above its recent baseline:
$$\text{Rising} = \mathbb{1}!\left[y_{t+1:t+2} \;\ge\; 1.2 \cdot \bar{y}_{\text{4wk}}\right]$$
We test four hypotheses — weather lag (H1), spatial spillover (H2), recurrence (H3), and green space (H4) — as engineered feature groups. To pick where limited effort goes first, we rank areas two ways:
$$R_{\mathrm{cases}} = p \times \hat{y}, \qquad R_{\mathrm{rate}} = \frac{p \times \hat{y}}{\mathrm{residents}} \times 10{,}000$$
And because missing a real outbreak is far worse than one extra reminder, we tune our threshold with a cost matrix that assumes a false negative costs five times a false positive:
$$\text{cost} = 5 \cdot \text{FN} + \text{FP}$$
The human layer. A forecast alone is just another dashboard. So every prediction ends in a ready-to-post message — what is rising, why (in plain words, never feature names), and one to three prevention steps that fit that area's likely cause this fortnight — handed to a community group-chat admin to share in the chats residents already use.
What we learned
- Prediction without action is a dashboard. The most important design decision was prescriptive: mapping each SHAP driver to an action with an owner and a timing. That is what turned a forecast into something a person can actually use.
- Honesty is a feature, not a caveat. Our training archive ends in November 2020, and the live feed only begins in September 2026. We learned to state exactly what is measured versus what is a target — we never claim a lead time until the backtest measures it, and we say "cluster activity," not "total cases," because that is what the data actually holds.
- Recall matters more than accuracy here. A missed rising area costs far more than an extra reminder, so we optimise for catching rises — measured through capture rate at K:
$$\text{capture}K = \frac{\sum{a \in \text{top}_K} c_a}{\sum_a c_a}$$
- The map is not the territory. Planning areas, Town Council boundaries, and MCST estates do not align, and no open dataset maps one to the other. Rather than imply a routing capability we could not support, we made every community action a checklist any local actor can pick up.
Challenges we faced
- The data does not want to be predictive. NEA's public cluster feed keeps no history — it is today's snapshot only. Our first act was a scheduled job that archives it daily into Delta, building the history the feed never kept, and doubling as our out-of-time test set.
- A five-year data gap. The training archives end in 2020; live data starts in 2026. We treat every daily snapshot as a prospective, out-of-time check, and a missed scoring day honestly lowers our reported coverage instead of being silently dropped.
- Building lean on a free platform. Databricks Free Edition is quota-limited and serverless. We pre-aggregate weather to daily before upload, keep workloads tiny, and design for batch inference rather than live serving.
- Saying no to what we could not prove. We resisted the tempting claim that community messaging reduces dengue — the evidence is weak — and promised a delivery mechanism, not a behaviour change. We also rejected routing to named institutions we could not map by boundary.
In the end, DengueRadar is our answer to a simple question: what if Singapore's residents knew two weeks earlier, with the reason and the remedy in hand — instead of finding out after the cluster already formed?
Built With
- data.gov.sg
- databricks
- mlflow
- sql
Log in or sign up for Devpost to join the conversation.