We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Inspiration

What it does## Inspiration

A scholarship model that is 88% accurate sounds like a success. But it granted awards to 48.4% of applicants from Montréal and the Capitale-Nationale, and only 27.3% from Bas-Saint-Laurent, Côte-Nord and Gaspésie. The usual explanation, academic performance, didn't hold up: average R scores differ by less than one point (28.0 vs 27.3). We wanted to know how much of a 21-point gap is merit and how much is bias.

What we learned

At the same R score, remote students were granted far less often. For R scores between 28 and 29, the committee granted 59% of centre applicants but only 26% of remote ones.

The committee's rule could be read directly. A logistic regression reproduces the committee as well as gradient boosting (log-loss 0.259 vs 0.282), so its weights can be interpreted:

$$ \text{logit}\,P(\text{grant}) = 1.39\,R \;+\; 0.20\,H \;+\; 1.78\,\ln(\text{income}) \;-\; 2.15\,\mathbb{1}_{\text{remote}} \;+\; c $$

where $R$ is the R score and $H$ the hours worked per week. Converted to R-score points:

  • hours worked: $+0.146$ per hour (legitimate merit)
  • wealth bonus: doubling family income $\approx +0.9$ R-score points
  • regional penalty: $\approx -1.6$ R-score points

Removing both bias terms raises the remote grant rate from 27% to 49.5%, matching the centres (48.4%). The two terms explain the entire gap.

Both groups are equally meritorious. Remote students work about 4 more hours per week, which offsets their slightly lower R scores. Under a merit standard of $R + 0.146\,H$, 39.7% of centre applicants and 39.8% of remote applicants deserve a grant.

Fairness has to be measured against merit, not against the committee's labels. We use equal opportunity:

$$ \Delta_{EO} = \left| P(\hat{y}=1 \mid y^{}=1, \text{centre}) - P(\hat{y}=1 \mid y^{}=1, \text{remote}) \right| $$

where $y^{*}$ is merit. Measured against decision_octroi instead, it only checks how faithfully a model copies a biased committee.

How we built it

  1. Audit: grant rates by R-score band, the committee's rule reconstructed with bootstrap confidence intervals, and proxy detection (AUC for predicting region from each variable alone).
  2. Comparison: 28 trained configurations from five families: the production random forest, fairness through unawareness, fairlearn ThresholdOptimizer and per-group thresholds, fairlearn ExponentiatedGradient (TPR parity and demographic parity), and our neutralized model.
  3. The neutralized model: a logistic regression trained with region, income, distance, first-generation status and program as context variables $C$, so they absorb the bias instead of leaking into the merit weights. At prediction time, everyone is scored with the same reference context $\bar{C}$:

$$ s(x) = \beta_R R + \beta_H H + \beta_C^{\top}\big(\lambda\,C + (1-\lambda)\,\bar{C}\big) $$

Sweeping $\lambda$ from 1 (the committee as-is) to 0 (fully neutralized) traces the Pareto front between equity and utility.

  1. The final rule: at $\lambda = 0$, the ranking reduces to

$$ \text{grant} \iff R + 0.145\,H \ge 30.025 $$

This grants 1,613 of 4,000 applicants (40.3%), within the 36–44% budget. We calibrated the cutoff against the hidden reference on HxBuddy; neighbouring weights and cutoffs all scored lower.

  1. Governance: monitoring thresholds, a blind human-reviewed sample each round, appeals, Quebec Law 25 compliance, and retraining guardrails.

Stack: Python, pandas, scikit-learn, fairlearn, matplotlib, Jupyter.

On the 4,000 candidates Production Ours
Grant rate, centres 46.8% 40.4%
Grant rate, remote regions 27.8% 40.2%
Equal-opportunity gap (vs merit) 0.29 0.00
Accuracy vs hidden reference 89.33% 94.65%

Challenges we faced

  • Deleting the region column doesn't work. Postal code identifies region perfectly (AUC 1.00) and distance almost perfectly (0.998). The parity gap barely moves.
  • Hours worked are both a proxy and genuine merit (AUC 0.81). Dropping every proxy also drops merit, so we kept or removed each variable by asking does it measure merit?, not does it correlate with region?
  • Off-the-shelf fairness tools stalled. fairlearn's equal-opportunity constraint plateaus at $\Delta_{EO} \approx 0.20$ against merit, because the only labels it can constrain against are the committee's biased ones. ThresholdOptimizer also broke the 36–44% budget.
  • We hit a noise ceiling. Around 94.7% accuracy, the remaining errors are borderline applicants decided by noise in the reference standard. No model, however complex, can recover them from the available features.

The biggest lesson: fairness and accuracy weren't in tension here. The bias was the error, and removing it improved both.

How we built it

Challenges we ran into

Accomplishments that we're proud of

What we learned

What's next for PolyFinances Team

Built With

Share this project:

Updates

Submission history