Inspiration
Buying an HDB flat is one of the biggest decisions most Singaporean households will ever make, and it's a decades-long commitment. Yet most buyers decide based on asking prices, agent opinions, and gut feel. Price data lives on data.gov.sg, location data lives on OneMap, and BTO launch information lives on HDB's site. Nobody puts them side by side.
We wanted to build something neutral: a tool that helps buyers track and compare resale flats and BTO launches with real data. It is deliberately not a financial advisor and never tells anyone what they can or can't afford.
What it does
- Resale Price Estimator. Every flat shows a data-backed price estimate from a model trained on public HDB resale transactions.
- Locality Score. Each flat is scored on how close it is to MRT stations, schools, supermarkets, and parks, with nearby amenities shown on a Leaflet map using OneMap tiles.
- Personalised weighting. In Settings, users rank the four amenity types by importance instead of setting numeric sliders, and the score reweights to match.
- BTO Tracker. A table of real past BTO exercises, plus one clearly labelled simulated "open" exercise to demo the live state.
How we built it
Price model. We trained a linear regression with scikit-learn on the data.gov.sg HDB resale dataset (~170,000 transactions). The inputs are town, flat type, floor area, storey range, and remaining lease. We predict the log of the price rather than the raw price:
\( \log(\text{price}) = \beta_0 + \beta_1 x_1 + \dots + \beta_n x_n \)
The model is trained once, offline, in train_model.py, and predictions are stored in SQLite. Nothing is computed live while the app runs.
Locality Score. For each flat, we compute the straight-line (haversine) distance to the nearest amenity in each category. The user's ranking is converted to fixed weights of 40/30/20/10%:
$$ \text{Locality Score} = \sum_{i=1}^{4} w_i \, s_i, \quad w = (0.4,\ 0.3,\ 0.2,\ 0.1) $$
Here \( s_i \) is the proximity score for the amenity the user ranked \( i \)-th.
App. Built with Flask and Jinja2 templates, styled with Bootstrap and a little vanilla JS, and backed by a single SQLite file. The stack is deliberately simple so it runs entirely on one machine with no build step and no cloud dependencies.
Challenges we ran into
- Funnel-shaped errors. Regressing on raw price produced errors that grew with price. Switching the target to
log(resale_price)fixed this. - Honest evaluation. If the 20 demo flats were in the training data, the model would just recall prices it had already seen. We held them out entirely, so each price shown is a genuine blind estimate.
- Distance to MRT as a price feature. Doing it properly meant geocoding all ~170,000 training rows, which wasn't feasible in our timeline. We dropped it from the price model, and location is captured by the Locality Score instead.
- API rate limits. All OneMap calls happen once, in an offline seed script. Results are cached to a committed JSON file, so the live app never calls OneMap directly.
- No live BTO launch to demo. Real launches don't run on our schedule. We backfilled past exercises from public information and added one simulated exercise that is clearly labelled as such.
- Working in parallel. We split the property detail page into a price section and a locality section, so two teammates could build it without editing the same lines.
Accomplishments that we're proud of
- Every price on the platform is a genuine blind estimate, tested against real held-out sales rather than recalled from training data.
- A Locality Score that reflects each user's own priorities, through a ranking interface that's simpler than sliders.
- Covering both buying pathways, resale and BTO, in one place.
- A demo that is fully self-contained and reliable: no live API calls, no cloud dependencies, and no risk of rate limits on presentation day.
- Being clear about what the tool is and isn't. We chose not to give affordability advice.
What we learned
- Simple models are easier to explain to users than complex ones, as long as you evaluate them honestly.
- Caching external API calls offline makes a demo far more reliable.
- Deciding what to leave out kept the project finishable. We cut live scraping, affordability advice, and a JS framework.
- Small UX choices matter. Ranking amenities is much more intuitive than setting percentage sliders.
What's next for KPC HDB Support system
We plan to move the pipeline onto Databricks:
- Lakeflow to ingest the source data into Delta tables
- Unity Catalog for governance of those tables
- MLflow to track the price model
- Databricks Apps to serve the user interface
We'd also like to expand beyond our 20 curated demo properties.
Log in or sign up for Devpost to join the conversation.