Inspiration

The brief gave us a scholarship committee that grants 48.4 % of applications from Montréal and Québec City but only 27.3 % from Bas-Saint-Laurent, Côte-Nord and Gaspésie. The average R score gap between those groups is only 0.66 point. At the same R score (28–30), a student from a big city got the scholarship 70.9 % of the time and a student from a remote region 40.8 %. I wanted to know how much of that gap was merit and how much was where people live, and then fix the part that wasn't merit. I was doing this alone against teams of four, so I knew I had to keep the solution simple enough to defend line by line.

What I learned

The first thing I did was rebuild the committee. A plain logistic regression explained its decisions as well as any bigger model I tried, and the regional term came out flat:

$$ \operatorname{logit} P(\text{grant}) = \beta^\top x - 1.90 \cdot \mathbb{1}[\text{remote}] $$

The penalty's 95 % confidence interval is [−2.07, −1.73]. A test of equal slopes between the two groups gave \( p = 0.82 \), so it wasn't a different rule per region, just a penalty added on top.

I thought dropping the region column would be enough. It wasn't. Postal code predicts the region perfectly (AUC 1.0), distance almost perfectly (0.998), and even hours worked (0.805) and family income (0.69) carry it. Income was the surprise: the committee rewards it, and income is lower in remote regions, so that reward was a hidden penalty worth about 29 % of the explicit one.

For fairness I chose equal opportunity instead of parity. The R score really is a bit lower in remote regions, and that justifies roughly 7 points of rate difference. What it doesn't justify is a different chance at the same merit:

$$ \text{EO gap} = P(\hat{y}=1 \mid \text{qualified, centre}) - P(\hat{y}=1 \mid \text{qualified, remote}) $$

How I built it

The model only learns from the historical decisions. The base score is the consensus rule

$$ s = z(R) + 0.185 \cdot z(\text{hours}) + 0.025 \cdot z(\log \text{income}) $$

with the regional penalty removed. On top of that, I added a residual from four TabM networks trained on the history on a Kaggle GPU, scaled by 2.5. The top 39.94 % (1,598 of 4,000) get the scholarship. To define what "qualified" means, I wrote five reviewer perspectives (merit, need, legal, data process, regional). I fixed the direction of each criterion myself and let the data set the size.

After the model come two juries and a harness. The juries vote on every borderline case and record their reasoning. The harness has four guard layers and an output check, and it can shift the cutoff, revert the jury or block publication. For the Pareto front, I swept the share of the regional penalty removed from 0 to 100 %. Agreement with the five references went up as the gap went down: from 0.223 to 0.019 on the gap, and from 0.925 to 0.976 on agreement.

Challenges

The jury was the hardest part to be honest about. I built it to swap decisions on borderline cases, and when I let it act, it proposed 26 swaps that made the equal-opportunity gap worse by 0.018, above my own 0.010 alert threshold. The harness pulled it out automatically. So the jury still runs on every batch, but only to audit. I was disappointed at first, but it ended up being the best proof that the guards do something.

I also had to stay strict about using the history only and declaring every setting. The 2.5 blend is a good example: it isn't on the Pareto front (0.5 is slightly better there), and I chose it for its historical log loss. I'd rather say that than hide it.

In the end, centres get 40.1 % and remote regions 39.7 %, with an impact ratio of 0.924 and 1,598 grants inside the budget.

Built With

Share this project:

Updates

Submission history