Inspiration
Track 2's problem is clean as far as ML tasks go: rank each user's own feed exposures by long_view likelihood, scored by
$$\text{primary} = \text{mean}(\text{GAUC}, \text{nDCG@5})$$
The starter kit came with a working FM baseline (0.5946 on hidden test) and a list of directions the organizers said nobody had tried yet.
What we built
The final model blends three things trained separately: an FM ranker with pairwise BPR loss (a 5-seed ensemble), a LightGBM ranker (num_leaves=2, linear_tree=True, lambdarank objective, 52% of the blend) using a feature set with a 5-day watch-depth decay and a native upload_type field, and an XGBoost ranker on the same causal features left un-bucketed (38%). FM takes the remaining 10%.
Validation primary 0.69943440. Test primary 0.68432260. Baseline is 0.5946, so that's +0.08972 absolute, +15.09% relative. python3 make_submission.py reproduces the whole thing from raw data in a couple of minutes on one CPU core.
How we built it
We ran this as an orchestrator plus agent loop, split across four tracks. On the main track, every round picked a direction from the starter kit's own untried list, implemented it, checked any hand-derived gradient against finite differences before trusting a real run on it, reproduced a known result exactly before building anything new on top of it, selected on validation only, and confirmed anything above 0.001 across five seeds before it counted. We set a convergence rule before starting instead of calling it on a hunch: three rounds in a row under 0.002 improvement and we stop. That happened after round 12. Switching the loss from pointwise to pairwise BPR was the biggest single jump; recency-decay features and activity-weighted sampling came after that.
We kept the loop running past convergence anyway, since there were still starter-kit directions we hadn't touched yet. Two earlier LightGBM and CatBoost attempts had been rejected for underperforming FM. It turned out the actual problem was the feature encoding: both had been forced through FM's bucketed representation, which throws away the continuous signal a tree model wants to see. Giving the GBM its own un-bucketed version of the same features closed the gap and then flipped it, and a later sweep found the score kept climbing as we shrank tree capacity all the way down to num_leaves=2. A jump that big and that smooth made us suspicious. Before promoting it we went looking specifically for a stable-sort artifact in the evaluation code and a leaked feature. Both checks came back clean, the result held across five seeds and a date-shifted split, so we kept it.
A teammate, Yixi, was working the same problem on a separate branch. She reused our FM training code as-is but added XGBoost as a third model and did her own feature and objective tuning on top. Her branch landed on the same 10/52/38 blend we ended up submitting, ahead of anything the main track had found by itself. Rather than take her numbers on trust, we retrained all three of her models from the raw CSVs, seed 0, none of her cached artifacts, and matched her result to eight decimal places before promoting it as the new best.
After that, a nine-round merge track went looking for anything left on the table by combining both lines of work, and a second teammate, Xuxia, stress-tested the merged blend from a third angle: per-segment weighting, rank fusion, calibrated fusion, multi-task stacking on top of it.
After this, we were finally able to converge onto our chosen model.
Challenges we ran into
Adding a fourth model to the blend raised validation by 0.00038, steady across five seeds, but test moved the wrong direction by 0.00065. We checked it two separate ways rather than just picking a side. Reselecting the blend weight under five-fold cross-validation landed on the exact same weight, which ruled out a grid-search overfit. A separate out-of-fold stacked meta-learner, which can't overfit a single split the way a grid search can, still scored below the plain three-model blend. Both checks pointed the same way: a real gap between the validation window and the test window, nothing wrong with how we'd searched. We dropped the fourth model anyway and kept the blend that had already been checked twice over.
There was also isotonic calibration, and we tried recalibrating each model's scores before blending, hoping it would smooth over the same inconsistency, and instead it crushed the tree models' scores from hundreds of thousands of distinct values down to a few dozen buckets. Validation fell by 0.061. Xuxia ran into the same wall independently, days later, on a different blend configuration: about 123,000 unique GBM scores collapsed to 37, validation down 0.036.
We also retested three decay-rate features, author popularity, duration bucket, hour of day, twice each: once against a large feature set, once against a stripped-down one, on the idea that a weak feature might win a split if it had less competition.
Built With
- lightgbm
- xgboost
Log in or sign up for Devpost to join the conversation.