Inspiration
Singapore is getting hotter. NEA recorded 29 days of high heat stress in 2025 (WBGT at 33°C or above), and has warned that extreme heat could arrive in March to May 2027 as El Niño weakens. The people most at risk are often the least visible: about 87,000 residents aged 65 and above lived alone in 2024, and 65,000+ seniors live in HDB 1 and 2-room flats. Heat harms them quietly at home, so it is rarely recorded as heat illness.
Active Ageing Centres (AACs) already know their vulnerable seniors. What they lack is a forward signal. NEA tells you where it is hot right now; planning tools model heat years ahead. Nothing tells a centre which days to step up outreach and which blocks to reach first. Even the government's new heat office says alerts alone won't reach everyone and that community touchpoints must check on seniors.
The real constraint is capacity. On a hot day, a centre might manage around 100 calls (our assumption, to verify with a centre) but serves 2,500+ seniors. So the question we set out to answer was simple: if a centre can only make so many calls, how does it make them count?
What it does
adem turns tomorrow's heat forecast into today's call list. It works in four steps:
- Assess (day before, 6pm): forecast tomorrow's heat stress, hour by hour, for every NEA station.
- Detect (day before, 6pm): rank HDB blocks by how many seniors will be exposed.
- Escalate (day before, evening): send each centre a ranked call list sized to what it can actually reach, with call scripts in English, Chinese, Malay and Tamil and the nearest cooling centre for each block.
- Mobilise (heat day, 9 to 11am): staff review the list and volunteers call before the heat peaks.
The split of responsibility is deliberate: adem decides days and blocks; AAC staff decide residents and follow-up. adem stores no names or phone numbers. It ranks blocks, and each centre matches them to its own member list using local knowledge. Staff can override any block, and each override is logged with a reason, so the ranking learns from the people on the ground.
Each block's priority is transparent:
$$ \text{Block priority} = \text{heat band} \times \text{senior exposure} \times (1 + \text{refuge gap}) \times \text{haze uplift} $$
- Heat band: tomorrow's forecast WBGT on NEA's official scale (Low below 31°C, Moderate 31 to 33°C, High from 33°C).
- Senior exposure: seniors 65+, small rental flats and older blocks. This is an area proxy, not a measure of individual vulnerability.
- Refuge gap: distance to the nearest air-conditioned cooling centre.
- Haze uplift: rises when PM2.5 is high.
Calls start at Moderate for the top-ranked blocks, and the full list goes out at High, to avoid alert fatigue.
How we built it
adem is designed as one Databricks loop from sensor to doorstep, using only open government data:
- Sources: live NEA APIs from data.gov.sg (WBGT from 30 stations, air temperature from 18 stations, humidity, rainfall, PM2.5 and NEA's 24-hour forecast), plus HDB Property Information, SingStat residents by subzone, age and dwelling type, URA Master Plan subzones, and AAC locations, geocoded with OneMap.
- Ingestion: Lakeflow Jobs poll the live APIs every 15 minutes and backfill history by date; Auto Loader brings static files into Delta.
- Transform: Lakeflow Declarative Pipelines take data through Bronze (raw), Silver (cleaned and aligned to 15-minute windows) and Gold (features), with data quality checks.
- Model: we forecast each station's hourly WBGT for tomorrow. A "tomorrow = today" baseline and LightGBM set the bar. A Temporal Fusion Transformer is kept only if it beats both on the same time split, on peak WBGT, peak timing and High-period recall. Everything is tracked in MLflow. Block heat is interpolated from nearby stations.
- Act: a Databricks App shows each centre its ranked call list, AI/BI dashboards and Genie give an island-wide view, and SQL Alerts flag live band changes during the day.
- Learn: staff reviews and overrides are stored in Lakebase and feed back into the ranking.
- Govern: Unity Catalog handles lineage, access control and tags, with no personal data anywhere in the pipeline.
Before committing to this design, we verified the data ourselves. We checked the live WBGT feed (30 stations reporting), confirmed that past NEA forecasts can be retrieved exactly as issued, and computed the 65,000+ figure directly from SingStat's subzone file.
Challenges we ran into
- The obvious datasets didn't work. The listed elderly population data turned out to be survey percentages, not block-level counts, and the senior centre dataset held national totals rather than locations. We rebuilt the data plan around SingStat subzone data, HDB block data and OneMap geocoding.
- Area data can't identify individuals. Subzone age profiles can't tell us who lives alone, who lacks air-conditioning, or how hot it is inside a flat. Instead of overclaiming, we made this the design: adem points to blocks and days, and staff choose residents.
- Heat between stations is an estimate. We plan to hold out each station in turn and report the interpolation error in °C, and to flag blocks that sit close to a band threshold.
- Indoor heat isn't measured. Sensors are outdoors, and homes can be hotter. We use block age and flat size as proxies and let staff add what they know.
- Keeping the backtest honest. A replay of a hot week must use only what was known at 6pm the day before. We use NEA's forecasts as issued and hold out a 2026 hot week the model has never seen.
- Short history. WBGT data only goes back to February 2025, and many stations are newer still. That is exactly why the Transformer has to earn its place against simpler baselines.
Accomplishments that we're proud of
- A tool that ends in an action. adem produces a call list volunteers can work through, not a map to interpret.
- Honesty built into the design. Limitations are stated on the poster together with what we do about each one, and our tests separate what a prototype can prove from what needs a real pilot.
- People stay in charge. Staff review every list, can override it, and unanswered calls follow the centre's own process rather than an automatic escalation.
- A grounded data plan. Every dataset is named, open and checked, and every key statistic is sourced.
What we learned
- Prioritisation has to be justified. A ranked list only matters because capacity is limited, so we had to show the gap between the calls a centre can make and the seniors it serves.
- Simple baselines first. A complex model is only worth it if it beats "tomorrow = today" on the measures that matter to the people using it.
- A completed call is not proof the score was right. Real impact runs through a chain (ranking, contact, help given, less exposure), and only the first links can be tested in two weeks.
- Check the data before designing around it. Several "obvious" datasets fell apart on inspection, and the live station counts had changed since the brief was written.
What's next for adem
- Two-week sprint: build the core loop for one AAC catchment, from forecast to ranked blocks to a staff-reviewed call list, and replay a held-out 2026 hot week.
- Test with real users: walk AAC staff or befriending volunteers through the list, compare our ranking blind against a simple heat plus elderly-density list, and record how often and why they override it.
- Pilot before the 2027 hot season: run with a centre ahead of March to May 2027, and measure whether top-priority blocks are reached by 11am, whether contacted seniors get cooling help or a referral when needed, and how many staff hours each senior reached costs.
- Stretch goals: learn from call outcomes, plan where new cooling centres would protect the most seniors, and link to the 300+ cooling centres that open when a heatwave is declared.
Built With
- databricks
- databricks-apps
- genie
- lakeflow
- lightgbm
- mlflow
- temporal-fusion-transformer
- unity-catalog
Log in or sign up for Devpost to join the conversation.