Inspiration
A student-aid model granted awards to 48.4% of applicants from Montréal and Québec City, but only 27.3% from remote regions. Grades explained only part of that gap. We wanted to know where the rest came from, and whether a fair model could stay accurate.
What it does
It ranks applicants on merit: academic performance (R score) and effort (hours worked during studies). Income and region can no longer push a score up or down. The top 40% of the cohort is funded, inside the fixed budget.
How we built it
- Audit: we measured the bias and hunted for proxies. Distance, hours and income all carry the region, so simply deleting the region column changed almost nothing.
- Model: a spline logistic model of the historical committee, then a counterfactual score. Each applicant keeps their own merit features, and everything else is replaced by shared reference profiles:
$$\text{score}(i) = \frac{1}{256}\sum_{p=1}^{256} P(\text{grant} \mid \text{R score}_i, \text{hours}_i, \text{profile}_p)$$
- Stability: averaged over 30 bootstrap seeds. 97% of applicants get the same decision from every seed.
- Result: grant rate of 40.0% in both remote and central regions (it was 27.8% vs 48.4%), with about 95% agreement on the indicative leaderboard.
Challenges
- The reference labels are hidden and differ from the committee's decisions, so we were predicting a target we never saw.
- Proxies everywhere: distance alone identifies the region with an AUC of 0.998.
- A noisy leaderboard: on 4,000 rows, many of our variants were within noise of each other.
What we learned
- "Fairness through unawareness" fails. Removing a sensitive column does not remove its signal.
- Defining merit is a value choice, not a statistical result, and it has to be stated openly.
- A leaderboard that shows exact scores leaks information about individual labels. We kept that exploration documented separately from the model, because it fits the evaluation set, not future applicants.
Built With
- jupyter
- matplotlib
- numpy
- pandas
- python
- scikit-learn
- scipy
Log in or sign up for Devpost to join the conversation.