Inspiration

Recommendation systems can look better on one visible score because of time leakage, seed luck, or one unusually easy user group. We built TraceRank around a stricter idea: evidence before promotion.

What it does

TraceRank is an autonomous machine-learning research system for ranking short-video candidates for each user. It begins with the organizer Factorization Machine baseline, declares one bounded hypothesis at a time, runs the experiment under fixed attempt and time limits, and records every command, metric, resource cost, failure, and keep/reject decision.

For KuaiRand-Pure, the selected candidate combines user and video context with attention over each user's last 20 positive long-view video and tag events. Six independently trained causal models vote by percentile rank inside each user's candidate list. This reduces seed noise and aligns the consensus with user-grouped ranking metrics.

For the separate KuaiRand-1K benchmark, the agent discovered that a simpler content-aware sparse Factorization Machine transferred better than the history model, so it retained the simpler candidate instead of forcing one architecture everywhere.

How we built it

  • Preserved the organizer loader, evaluator, baseline, and submission checker.
  • Used chronological training, validation, and forward-time screens.
  • Checked low-, medium-, and high-activity users plus early and late date slices.
  • Required repeatable evidence across fixed seeds before promotion.
  • Stored append-only JSONL experiment ledgers, SHA-256 manifests, exact commands, runtimes, and memory peaks.
  • Built deterministic acquisition, privacy, release, and label-boundary checks.
  • Adapted disciplined experiment-loop ideas from Karpathy autoresearch, AIDE, and FML-Bench while keeping their workloads separate from our solution.

Challenges

The data contains feedback only for videos users were actually shown, so it is not a random view of all possible recommendations. User histories are incomplete, positive outcomes are sparse, and temporal drift can make a model look strong on one period but weak later. The published materials also contain a metric contradiction, so we treated the executable organizer evaluator, GAUC and nDCG@5, as authoritative while documenting the discrepancy.

Many seemingly promising ideas failed: listwise losses, deeper networks, captions, richer categories, multiple action histories, time features, residual rankers, hard target matching, and larger ensembles. TraceRank rejected them when gains failed chronological or subgroup gates.

Accomplishments

On local validation:

  • KuaiRand-Pure: 0.605375 primary, GAUC 0.672521, and nDCG@5 0.538228, versus the published 0.601600 baseline.
  • KuaiRand-1K: 0.653747 primary, GAUC 0.688786, and nDCG@5 0.618707, versus our fixed 0.644227 local base.
  • The final Pure prediction file passed a label-blind 170,588-row alignment check.
  • The complete research trail remains auditable and reproducible on one Apple-silicon laptop with zero CUDA GPU-hours.

These are local validation results, not leaderboard or final-test scores. The final-test outcome remains unknown.

What we learned

More complexity is not automatically better. Recent meaningful viewing history helped the required Pure task, while a smaller content model transferred better on 1K. The most valuable part of an autonomous agent is not how many experiments it runs; it is whether it knows when evidence is too weak to justify promotion.

What's next

We would extend the same evidence gates to online evaluation, richer privacy-preserving histories, and distribution-shift monitoring. The immediate artifact is a reproducible, judge-auditable ranking research system with explicit limitations rather than an inflated performance claim.

Built With

Share this project:

Updates

Submission history