Inspiration
A committee awards scholarships, and a model trained on its past decisions is used in production. At equal academic merit, applicants from remote regions are funded far less often than applicants from Montréal and the Capitale-Nationale: 27.3 % against 48.4 %. We wanted to find where that gap comes from before trying to fix it.
What it does
It diagnoses the bias, corrects it under the same budget, and proposes a monitoring plan.
- Diagnosis: two channels. A uniform regional penalty worth about 1.6 R-score points, which explains about 18 of the 21 points of gap. And a reward for household income, which is lower in remote regions.
- Correction. Decide on R score plus hours worked only, with the committee's own weights, and fund the top 40 %.
- Result. The grant rate is 40.0 % in both groups, and the estimated equal-opportunity gap goes from 0.27 to about 0. The indicative accuracy on HxBuddy is 94.63 %, against 89.33 % for the production model.
How we built it
- A logistic model of the committee with an explicit region term, so the penalty is isolated instead of leaking into proxies. It fits as well as gradient boosting (AUC 0.956 against 0.952).
- SHAP on the production model: geography pushes remote applicants down 4.6 times more than their real R-score difference does.
- A decision score restricted to R score and hours worked. Region is used to measure, never to decide.
- A comparison with Fairlearn (ExponentiatedGradient and ThresholdOptimizer) at the same budget, and a Pareto front of fairness against agreement.
- Python, pandas, statsmodels, scikit-learn, Fairlearn, SHAP. Claude (Anthropic) was used as a coding assistant.
Challenges we ran into
- Deleting the region column does not help. Distance predicts the region almost perfectly (AUC 0.998), and postal code, hours worked and income re-encode it too.
- The labels themselves are biased. A fairness constraint measured on biased labels cannot remove label bias: equal opportunity enforced on the committee's labels only moves the gap from 0.28 to 0.19.
- The reference standard is hidden. We tested six global hypotheses against the indicative scorer; this was not row-level probing. The reference also contains noise we cannot predict.
Accomplishments that we're proud of
- A rule anyone can read and recompute by hand: R score, plus a fixed weight per weekly hour worked.
- An explanation for each applicant, for example: "Refused: score 29.57 against a cut-off of 30.06. Region, income and distance were not used."
- Equal grant rates that follow from the de-biased score, not from a quota.
- A production monitoring plan with metrics, frequencies, alert thresholds and actions.
What we learned
- Find the mechanism of the bias first; the fix follows from it.
- Fairness metrics are only as good as the labels they are measured against.
- The value choices (neutralizing income, keeping hours worked) belong to the institution and must be stated openly.
What's next for ÉquiAlgo
- Have the ethics committee validate the weight of hours worked, which is also a regional proxy.
- Send decisions near the cut-off (about 9 % of files) to human review, with the score breakdown.
- Audit other sensitive attributes: only region is available in this data.
- Run the monitoring plan on each new cohort, with a blind re-assessment sample.
Built With
- fairlearn
- pandas
- python
- scikit-learn
- shap
- statsmodels
Log in or sign up for Devpost to join the conversation.