Inspiration

Singapore has invested heavily in reducing flood risk, but intense rainfall can still create highly localized flash-flood conditions. We wanted to build something that goes beyond showing where rain is falling and instead answers a more useful question:

"Where is flood risk building, how likely is flooding, and how much warning can we provide?"

This inspired FloodSense, an AI-powered flood intelligence system designed around Singapore's open environmental data.

What we built

FloodSense combines rainfall observations, flood-event information, and spatial risk data to transform raw observations into a localized flood-risk signal.

The system produces:

  • A real-time flood-risk map
  • Flood probability for individual locations
  • A 0–100 FloodSense Risk Score
  • Explainable risk drivers
  • Historical rainfall and flood-event analytics
  • FloodSense risk signals for rapidly increasing risk
  • A rainfall "What-If" simulator to explore how risk changes under higher rainfall intensity

The goal is not to replace official PUB warnings. FloodSense is designed as a decision-support layer that can help planners and response teams interpret multiple data sources and prioritize attention.

How we built it

The project was designed as an end-to-end data and AI pipeline on Databricks.

Public data is ingested into a Bronze layer, cleaned and transformed into Silver tables, and converted into Gold feature tables for machine learning.

We engineered causal temporal rainfall features including short- and long-term accumulation, rolling statistics, rainfall trends, rate of change, lag features, and antecedent rainfall conditions.

Multiple models were evaluated, with model selection based primarily on PR-AUC because flood events are highly imbalanced. MLflow is used to track experiments and model performance.

A validation-selected operating threshold is then applied to the frozen test set. This separates model selection from the final evaluation and prevents test-set leakage.

The final product combines the prediction engine with a FastAPI backend and an interactive Streamlit dashboard.

Challenges

One of our biggest challenges was the extreme imbalance between normal observations and actual flood events. A model can achieve very high accuracy simply by predicting that flooding will not happen, which would make it practically useless.

We therefore focused on PR-AUC, event recall, lead time, false-alarm rate, and threshold selection rather than treating accuracy as the only measure of success.

Another challenge was ensuring that the model remained causal. Features must only use information that would have been available at prediction time. We implemented automated leakage checks to verify that future observations cannot influence earlier predictions.

During development, direct access to some Singapore government APIs was restricted by the development environment. We therefore used synthetic data to validate the complete pipeline and evaluation framework while keeping the architecture ready for the official Singapore datasets. Synthetic validation results are explicitly labelled as such and are not presented as real-world Singapore predictive performance.

What we learned

We learned that building an AI system for disaster prediction is not just a modelling problem. Data quality, temporal alignment, event definition, class imbalance, threshold selection, and leakage prevention can matter as much as the choice of algorithm.

Most importantly, we learned that a 99% accurate flood model can still be a bad flood model if it misses the floods.

FloodSense therefore treats accuracy as an operating-point constraint while using event detection and warning lead time to measure practical usefulness.

Impact

FloodSense is intended to help urban planners, infrastructure teams, and emergency-response stakeholders move from reactive analysis toward earlier, data-driven flood-risk awareness.

The architecture is designed to scale from a prototype into a continuously updated system using Singapore's public data infrastructure.

Our next step is to evaluate and calibrate the complete pipeline on real Singapore observations and work toward deployment-grade reliability.

Built With

Share this project:

Updates

Submission history