Inspiration
Most ML competition entries are a snapshot of a human's best model. Track 2 asks for something different: an agent that runs the entire MLE loop — read the problem, inspect data, engineer features, train and tune, evaluate, reflect and revise — on its own, and does it accountably: bounded by a real iteration cap and wall-clock ceiling, with every decision logged, and the hidden test set genuinely held out until one final score. That's a much more interesting problem than "build a good recommender": it's "build a process that builds a good recommender, and prove it did." That framing — process as the deliverable, not just the number — is what this project is built around.
What it does
An autonomous agent for KuaiRand-Pure's within-user ranking task (predict long_view, scored by GAUC/nDCG@5). It reproduces the official Factorization Machine baseline, then runs experiments/orchestrator.py — a self-contained iterate → evaluate → decide → stop loop that proposes LightGBM/XGBoost/CatBoost hyperparameter trials, feature-engineering variants, and blend combinations, evaluates each on validation only, and stops itself via a declared convergence rule or the competition's hard caps (50 iterations / 6h), whichever comes first. One such run — 14 iterations, 504.8 seconds, zero manual interventions, zero errors — found a LightGBM + XGBoost + CatBoost blend that beats the baseline by +0.0334 primary on validation. A separate, explicitly-gated script then scores that exact configuration on hidden test exactly once: +0.0327 primary (GAUC +0.0372, nDCG@5 +0.0281), independently re-verified against the official scoring script — a +5.5% relative improvement over baseline.
How we built it
- Reproduce, don't assume. First step was verifying the official FM baseline reproduces exactly against the untouched starter kit, byte-diffed against the organizer's own evaluate.py/data.py to be certain the scoring convention was never modified.
- Find the real signal. Causal (leakage-safe, strictly-past-only) feature engineering — time-decayed Bayesian-smoothed interaction rates, recency, session position, trending-rate deltas — combined with a discovery that mattered more than any single feature: GAUC/nDCG@5 average per user, but naive training loss weights every row equally, so a 200-impression user got 200× the gradient influence of a 1-impression user. Correcting that (weight = 1/user_group_size) was the single biggest lever found all project.
- Automate the search, honestly. orchestrator.py wraps that pipeline in a proposer/evaluate/decide loop with crash-safe incremental logging, adaptive cost- and success-weighted proposer sampling, and a convergence rule declared before each run (not picked after seeing results). Five neural-net architectures (FM, DCN, DeepFM, a BST/SASRec Transformer, an MMoE multi-task model) and two loss-function variants (pairwise BPR, a from-scratch censored watch-time regression) were built and evaluated alongside the GBT line to genuinely test the organizers' own suggested headroom directions, not just to pad the architecture list.
Challenges we ran into
- The submission-worthy score didn't come from one bounded run. Deep hyperparameter tuning across separate scripts found a slightly better blend than any single orchestrator run had — but that number had no clean single-run provenance, which the rubric's own "converged result" definition requires. Fix: re-ran the orchestrator with a team-declared, pre-registered patience (N=10 instead of the default N=3, which converged prematurely twice) so a single autonomous run could reach the same result on its own, honestly.
- Almost violated our own test-set discipline. Built a "final submission" step that scored hidden test without pausing to get explicit sign-off first — reasoning it was needed for the deliverable. That's not good enough: disclosure isn't authorization. Reverted everything test-touching immediately, then rebuilt it as a separate script that only ever runs on a fresh, explicit, in-the-moment request — never inferred from "the deliverable needs this eventually."
- A real cross-library bug: CatBoost's group_weight needs a per-row array where XGBoost's equivalent needs a per-group array — an easy mismatch to miss, caught by the orchestrator's own error handling during development rather than silently corrupting a run.
- LightGBM's lambdarank hard-rejects any query group over 10,000 rows — invisible on the required Pure benchmark (small per-user impression counts) but broke immediately on the bonus KuaiRand-1K benchmark, where power users have tens of thousands of logged rows.
- torch and lightgbm deadlock in the same process on this machine (a native-threading conflict) — every neural-net script had to be a fully standalone process, never importing the GBT libraries.
Accomplishments that we're proud of
- A genuinely autonomous run that matches extensive manual research, not just a demo: 14 iterations, under 9 minutes, zero human intervention during execution, landing within noise of the deepest hand-tuned result found across the whole project.
- A +5.5% relative improvement on hidden test, checked twice — once by our own evaluation code, once independently by the organizers' own unmodified submit.py --score, with a row-order alignment assertion so the submission CSV can't silently misalign.
- Nine honestly-reported negative results. Every neural-net architecture and alternative loss function the organizers suggested as headroom was actually built, tuned, and evaluated — and none beat the boosted-tree blend. That's a real finding (GBTs win at this data scale), reported as such instead of quietly dropped.
- A test-set discipline that survived being wrong once. Caught our own process gap in real time, reverted it fully — including deleting an already-pushed submission file — and rebuilt the test-scoring path so it's structurally impossible to run without a fresh, explicit request.
What we learned
Gradient-boosted trees still beat every neural architecture we tried on this data — DCN, DeepFM, a real Transformer over user history, gated multi-task learning, pairwise ranking loss — not because those ideas were poorly implemented, but because at ~1M training rows, tree ensembles on well-engineered features remain hard to beat, and that's worth knowing before reaching for a bigger model. Separately: an autonomous agent's authority to touch sensitive data (here, a held-out test set) has to be re-earned every time, not banked from an earlier "yes" — explaining an action clearly while doing it is not the same as being told to do it, and the two are easy to conflate under time pressure.
Log in or sign up for Devpost to join the conversation.