Inspiration and problem

Recommendation experiments are easy to overfit to one seed, accidentally leak test information, or leave the repository in a broken state. This project asks an LLM coding agent to improve the official KuaiRand-Pure FM baseline while constraining it to the organizer's fixed date split and frozen evaluator. The goal is such that every hypothesis, code change, result, error, and recovery must remain auditable.

What it does

The agent reads the editable pipeline source and previous run history, then proposes one hypothesis and an exact code diff. The harness validates a strict JSON schema and file allowlist, applies the proposal only inside a scratch workspace, compiles it, and evaluates it over three deterministic seeds. A proposal is kept only if it clears a two-sigma noise threshold with a minimum absolute improvement of 0.001; otherwise all changed files are restored.

The successful proposal replaced pointwise log loss with a within-user pairwise ranking objective and changed batching so each user's impressions stay together. This directly aligns training with within-user GAUC and nDCG@5. The kept result improved test primary from the official 0.5946 baseline to a three-seed mean of 0.596750, an absolute gain of 0.002150.

The final recommender and submission generation are offline: they use Python and NumPy only. An OpenAI model was used as the coding agent during development; local Ollama and MLX models were also benchmarked, but their proposals were not used for the final score because they did not reliably reach evaluation.

How it addresses the challenge

  • Uses the fixed KuaiRand-Pure train/validation/test dates and native long_view label.
  • Reports the frozen GAUC and nDCG@5 metrics and their mean.
  • Selects experiments using validation performance and preserves the exact submission row order.
  • Logs hypotheses, diffs, metrics, errors, recoveries, token usage, wall time, and manual interventions for every iteration.
  • Runs proposed code in an isolated workspace and automatically reverts syntax, runtime, timeout, and statistically insignificant failures.

Development tools

  • Visual Studio Code and its integrated terminal
  • Git and GitHub
  • macOS on Apple silicon

APIs

  • OpenAI API (gpt-5) for the canonical research-agent run
  • Ollama's local HTTP API for offline model experiments
  • No API is required to train the final FM or generate its submission output

Libraries and frameworks

  • Python standard library
  • NumPy for the FM, training loop, checkpoint, and ranking scores
  • OpenAI Python SDK and python-dotenv for the canonical agent provider
  • Ollama for local inference experiments
  • MLX, MLX-VLM, Hugging Face model artifacts, and Jinja2 for the Gemma LoRA pilot

Dataset and assets

  • KuaiRand-Pure, using the organizer-pinned date split:
    • train: 2022-04-08 through 2022-04-21
    • validation: 2022-04-22 through 2022-04-28
    • test: 2022-04-29 through 2022-05-08
  • No external labels or manually created relevance judgments were used.
  • The dataset is intentionally excluded from Git and downloaded separately.

Results

On the canonical three-seed run, the validation-best model reached GAUC 0.668762, nDCG@5 0.536793, and primary 0.602778. On the held-out local test split it reached GAUC 0.663094, nDCG@5 0.530406, and primary 0.596750. See submission/RESULTS_AND_RESOURCES.md for the baseline deltas and full resource accounting.

Robustness and autonomy

The canonical log contains one kept proposal, three automatically reverted follow-up proposals, a baseline record, and one explicitly marked manual five-seed confirmation. The harness never silently patches a failed response: schema, replacement, syntax, subprocess, and metric failures are recorded with their recovery outcome. Convergence was first satisfied after logged iteration 4; the later confirmation and resumed experiment remain in the audit trail.

Challenges and lessons

The most important bug was conceptual rather than syntactic: row-randomized batches rarely contained multiple impressions from one user, so early pairwise objectives silently fell back to pointwise behavior. Introducing a user-grouped sampler made the intended objective real. We also learned that exact-text patch schemas are reliable for strong API models but brittle for small local models.

Limitations and future work

  • The measured gain is statistically verified but modest.
  • The canonical discovery used a hosted LLM API; the final recommender is offline, but local coding models did not yet match the API model's patch reliability.
  • Only KuaiRand-Pure was completed; the optional 1k and 27k benchmarks remain.
  • With more time, we would give local models a constrained experiment-action schema, collect more diverse verified examples, and evaluate several coding- specialized local checkpoints before another fine-tuning attempt.

Built With

Share this project:

Updates

Submission history