Inspiration

Volunteering as Silver Generation Ambassadors with the Silver Generation Office, we spent our time knocking on doors and checking in on seniors at home. Home visits work, but there are only so many volunteers and only so many doors we can knock on in a day. So which doors should we knock on first?

That's the gap we wanted to tackle. Singapore is ageing fast, with 18.8% of residents aged 65 and over in 2025. Yet, almost every public dataset about seniors counts people who come forward for help. The most isolated seniors are, by definition, the hardest to see in the data.

What it does

Matcha Latte(r) helps outreach teams decide where to look first. We rank places, not people: every subzone in Singapore is scored on three evidence-backed warning signs.

  • No one at home: estimated seniors living alone, with men living alone treated as a higher-risk subgroup
  • Housing strain: the share of seniors in HDB 1–2 room flats
  • Community access: everyday places within walking reach, and how stretched nearby AACs are

Instead of hiding uncertainty behind one score, we stress-test 10,000 plausible weightings and give each area a priority stability: the share of scenarios in which it stays in the top 10. Coordinators get a map and a card for each area with the reasons and a suggested action. They can also ask Genie plain-English questions like "Which areas stay high priority if housing weight is halved?" Results roll up from subzones to planning areas and towns.

How we built it

Round 1 is our design stage, so we focused on getting the foundations right before writing code.

  • Research first: we grounded each warning sign in Singapore studies, including the Singapore Chinese Health Study (16,943 older adults) and a 2024 Central Singapore study on housing and isolation.
  • Data due diligence: we opened every suggested A1 dataset and checked what it actually contains. Where one lacked neighbourhood-level detail, we found an official open-data substitute.
  • Model design: living-alone figures are estimated from subzone age and sex mix, national living-arrangement rates and planning-area household data. Community access uses a two-step floating catchment method. The three signs are combined with SMAA-2, which tests thousands of weightings instead of picking one.
  • Architecture: a bronze, silver and gold pipeline on Databricks, governed with Unity Catalog, tracked in MLflow, and surfaced through an AI/BI dashboard and Genie.

Challenges we ran into

  • There's no ground truth. Nobody measures isolation by neighbourhood, and the obvious labels (ComCare, AAC attendance, hospital use) only capture seniors who are already visible. Training a classifier on them would automate the very blind spot we're trying to fix. So we built a transparent ranking model and measured how much it depends on our assumptions.
  • Data isn't always what the title says. Some suggested datasets turned out to be national totals, and the public AAC location file dates from before centres became AACs. We had to track down subzone-level official sources and plan to compile current AAC locations ourselves.
  • Working within Free Edition. Outbound internet is restricted, so we planned to upload all files to Unity Catalog volumes and run geocoding locally. Model-serving availability is limited, so plain-language reasons come from templates, with AI-written reasons as a stretch goal.
  • Not overclaiming. It was tempting to say "94% chance of isolation". That would be wrong: priority stability measures confidence in the ranking, not the probability that anyone is isolated.

Accomplishments that we're proud of

  • Checking every statistic in our submission against its original source
  • Designing governance in from the start: coordinators see subzone results only, small counts are suppressed, and no individual data is used anywhere

What we learned

  • Isolation and loneliness are different things. And living alone, while a real risk factor, misses most of the picture: 85.6% of socially disconnected seniors in the largest Singapore study lived with others.
  • Sometimes the right ML decision is not to train a model. Without honest labels, a transparent and testable ranking beats a black box.
  • Statistics don't translate directly into weights. An odds ratio tells you direction and strength, not how many points to give a sign. Ranking by rates also gives different answers from ranking by counts, so we report both.
  • Read the data before trusting the label. Half our early surprises came from opening files and finding something different from what the title promised.

What's next for Matcha Latte(r)

  • The two-week build: week 1 for data and the three signs; week 2 for ranking, tests, the dashboard and the demo.
  • Learning from fieldwork: coordinators log visit outcomes, and the ranking learns from them, with about 15% of visits sent outside the top 10 to catch what the model misses.
  • An AAC placement recommender showing where a new centre would relieve the most pressure.
  • A 2030 view that projects where needs are heading as the population ages.
  • Partnering with AIC and the Silver Generation Office for centre capacity data and finer-grained tables, which would make the estimates sharper.

Built With

Share this project:

Updates

Submission history