Inspiration
What it does## Inspiration
A scholarship model that is 88% accurate sounds like a success. But it granted awards to 48.4% of applicants from Montréal and the Capitale-Nationale, and only 27.3% from Bas-Saint-Laurent, Côte-Nord and Gaspésie. The usual explanation, academic performance, didn't hold up: average R scores differ by less than one point (28.0 vs 27.3). We wanted to know how much of a 21-point gap is merit and how much is bias.
What we learned
At the same R score, remote students were granted far less often. For R scores between 28 and 29, the committee granted 59% of centre applicants but only 26% of remote ones.
The committee's rule could be read directly. A logistic regression reproduces the committee as well as gradient boosting (log-loss 0.259 vs 0.282), so its weights can be interpreted:
$$ \text{logit}\,P(\text{grant}) = 1.39\,R \;+\; 0.20\,H \;+\; 1.78\,\ln(\text{income}) \;-\; 2.15\,\mathbb{1}_{\text{remote}} \;+\; c $$
where $R$ is the R score and $H$ the hours worked per week. Converted to R-score points:
- hours worked: $+0.146$ per hour (legitimate merit)
- wealth bonus: doubling family income $\approx +0.9$ R-score points
- regional penalty: $\approx -1.6$ R-score points
Removing both bias terms raises the remote grant rate from 27% to 49.5%, matching the centres (48.4%). The two terms explain the entire gap.
Both groups are equally meritorious. Remote students work about 4 more hours per week, which offsets their slightly lower R scores. Under a merit standard of $R + 0.146\,H$, 39.7% of centre applicants and 39.8% of remote applicants deserve a grant.
Fairness has to be measured against merit, not against the committee's labels. We use equal opportunity:
$$ \Delta_{EO} = \left| P(\hat{y}=1 \mid y^{}=1, \text{centre}) - P(\hat{y}=1 \mid y^{}=1, \text{remote}) \right| $$
where $y^{*}$ is merit. Measured against decision_octroi instead, it only checks how faithfully a model copies a biased committee.
How we built it
- Audit: grant rates by R-score band, the committee's rule reconstructed with bootstrap confidence intervals, and proxy detection (AUC for predicting region from each variable alone).
- Comparison: 28 trained configurations from five families: the production random forest, fairness through unawareness, fairlearn
ThresholdOptimizerand per-group thresholds, fairlearnExponentiatedGradient(TPR parity and demographic parity), and our neutralized model. - The neutralized model: a logistic regression trained with region, income, distance, first-generation status and program as context variables $C$, so they absorb the bias instead of leaking into the merit weights. At prediction time, everyone is scored with the same reference context $\bar{C}$:
$$ s(x) = \beta_R R + \beta_H H + \beta_C^{\top}\big(\lambda\,C + (1-\lambda)\,\bar{C}\big) $$
Sweeping $\lambda$ from 1 (the committee as-is) to 0 (fully neutralized) traces the Pareto front between equity and utility.
- The final rule: at $\lambda = 0$, the ranking reduces to
$$ \text{grant} \iff R + 0.145\,H \ge 30.025 $$
This grants 1,613 of 4,000 applicants (40.3%), within the 36–44% budget. We calibrated the cutoff against the hidden reference on HxBuddy; neighbouring weights and cutoffs all scored lower.
- Governance: monitoring thresholds, a blind human-reviewed sample each round, appeals, Quebec Law 25 compliance, and retraining guardrails.
Stack: Python, pandas, scikit-learn, fairlearn, matplotlib, Jupyter.
| On the 4,000 candidates | Production | Ours |
|---|---|---|
| Grant rate, centres | 46.8% | 40.4% |
| Grant rate, remote regions | 27.8% | 40.2% |
| Equal-opportunity gap (vs merit) | 0.29 | 0.00 |
| Accuracy vs hidden reference | 89.33% | 94.65% |
Challenges we faced
- Deleting the region column doesn't work. Postal code identifies region perfectly (AUC 1.00) and distance almost perfectly (0.998). The parity gap barely moves.
- Hours worked are both a proxy and genuine merit (AUC 0.81). Dropping every proxy also drops merit, so we kept or removed each variable by asking does it measure merit?, not does it correlate with region?
- Off-the-shelf fairness tools stalled. fairlearn's equal-opportunity constraint plateaus at $\Delta_{EO} \approx 0.20$ against merit, because the only labels it can constrain against are the committee's biased ones.
ThresholdOptimizeralso broke the 36–44% budget. - We hit a noise ceiling. Around 94.7% accuracy, the remaining errors are borderline applicants decided by noise in the reference standard. No model, however complex, can recover them from the available features.
The biggest lesson: fairness and accuracy weren't in tension here. The bias was the error, and removing it improved both.
How we built it
Challenges we ran into
Accomplishments that we're proud of
What we learned
What's next for PolyFinances Team
Built With
- api
- python
Log in or sign up for Devpost to join the conversation.