For Agents Page - An Autonomous ML Research Agent for KuaiRand-Pure Ranking

Devpost written project description — TikTok TechJam, Track 2

How our solution addresses the problem statement

The track asks for an autonomous agent that improves a Factorisation Machine (FM) baseline on the KuaiRand-Pure ranking task, scored by primary = mean(GAUC, nDCG@5) on the validation split, and judged as much on reasoning quality, autonomy, and cost-efficiency as on raw accuracy gains.

Our answer is a three-agent loop — Research → Coding → Evaluator — wired together by a deterministic orchestrator, deliberately built so that every part of the control flow is fixed code, not an LLM's judgment call, and every part that genuinely needs reasoning gets a dedicated agent:

  • Research Agent (agent/research/) proposes one concrete, literature- grounded hypothesis per iteration. It uses a breadth-then-depth flow: generate several shallow candidate directions cheaply, filter out anything matching a known dead end or a prior attempt (via a persistent, cross-run "Do/Don't" ledger and a similarity check), then write one full structured proposal for the single best survivor. Every claim is grounded in a bundled citation catalog we verified against the actual papers (not model memory), and every proposal passes a safety scan that rejects anything trying to manipulate the benchmark or leak test-split information into its own reasoning.

  • Coding Agent (agent/coding/) turns one hypothesis into a runnable solution directory. Before ever spending a full multi-seed training run, it runs a cheap inner loop — generate → static checks → a small smoke run → repair — so a NameError never costs a real attempt. Each new idea builds on the current best checkpoint rather than always starting from the static baseline, so improvements compound instead of resetting every iteration. It's permitted to use any library actually installed in the environment (measured by import availability, not a hand-maintained allowlist), so long as it stays inside a documented sandbox (no network, no subprocess spawning, one CPU core, a fixed per-run timeout).

  • Evaluator Agent (agent/evaluator/) runs each candidate and decides ACCEPT / REVERT / ABANDON. It never trusts a solution's self-reported score. Every result is independently re-derived from the raw persisted predictions through the competition's own unmodified evaluate.py (agent/verification.py), so a generated train.py can never just claim a good number. Hidden test-split metrics are quarantined at the executor level and structurally blocked from ever reaching an agent-facing payload. This ensures that no agent ever sees them, at any point.

  • Orchestrator (agent/orchestrator.py, agent/convergence.py) owns the parts that don't need judgment: the competition's stopping rule (50 iterations, or 6h wall-clock, or 3 consecutive iterations without a validation-primary improvement greater than 0.002 — whichever fires first), a two-tier retry/abandonment policy (3 repair attempts per idea before abandoning it; escalate to a human only after two ideas in a row fail outright), and per-iteration token/cost accounting across all three agents.

A separate one-time exploration campaign (scripts/seed_findings.py, docs/exploration-campaign.md) runs the Starter Kit's own ranked list of "unexplored directions" (ranking loss, sequence modeling, multi-task heads, watch-time modeling, time drift, unbiased validation) through the real Coding → Executor → Evaluator loop before the graded run starts, so the Research Agent begins with an honestly measured Do/Don't ledger instead of an empty one or a set of assumptions.

Development tools used

  • Claude Code (Anthropic's CLI coding agent) — used throughout development; visible directly in the repo's own commit history via co-author trailers (Claude Sonnet 5, Claude Opus 5) across most of the project.
  • VS Code — primary editor.
  • Git / GitHub for version control and PR-based review between teammates.
  • pytest as the test runner during development (717 tests as of the pre-graded-run tag).

APIs used

  • OpenAI API (gpt-5, via the openai Python SDK) — the only runtime API our agent pipeline calls. All three LLM-backed agents (Research, Coding, Evaluator) share one client through dependency injection; the model is configurable via CODING_AGENT_MODEL. An --offline mode swaps in a deterministic, zero-cost template library and rule-based evaluator so the entire pipeline can also run end-to-end with no API key and no spend, for testing.

Libraries and frameworks used

Pinned in requirements.txt:

  • NumPy (>=1.24) — the baseline FM implementation and most harness code.
  • PyTorch (>=2.2) — for solutions whose hypothesis needs gradients (sequence models, multi-task heads, DeepFM-style architectures), so the Coding Agent doesn't have to hand-derive backprop.
  • openai (>=1.40) — Python SDK for the GPT-5 calls above.
  • PyYAML (>=6.0) — parses solution/config.yaml; a solution directory falls back to a small stdlib YAML reader when PyYAML isn't installed, and a test asserts the two agree.
  • pytest (>=8.0) — the test suite.

Also available to (and usable by) generated solutions, though not currently pinned in requirements.txt — pandas, scikit-learn, and scipy were installed on the development machine and are measured as available via importlib.util.find_spec rather than declared in a static allowlist. (Flagging this as a loose end: if the graded-run environment is provisioned strictly from requirements.txt, a solution the Coding Agent writes expecting these libraries would fail to reproduce there — worth pinning them explicitly before the graded run.)

Datasets and assets used

  • KuaiRand-Pure (via the competition's Starter Kit, originally from Zenodo) — the sequential recommendation/ranking dataset this track is built on: ~1.14M training rows, ~125K validation rows, ~171K test rows of logged interactions, plus a separate ~1.18M-row randomized-exposure log used only as a read-only diagnostic (never as training data).
  • The Starter Kit's vendored, unmodified scoring harness — evaluate.py (the sole scoring authority for GAUC / nDCG@5 / primary), data.py (official date-based train/valid/test splits and feature encoding), and baseline.py (the FM/pop/random baselines we measure against) — copied byte-for-byte with SHA-256 provenance tracking so we can prove none of it was modified.
  • A bundled literature citation catalog (agent/research/references.json) grounding the Research Agent's proposals — every entry checked against the actual paper text (e.g., Covington et al. 2016 on YouTube recommendations; Schnabel et al. 2016 on propensity-weighted evaluation) rather than trusted from model memory.
  • A persistent cross-run findings ledger (agent/research/findings.jsonl), seeded by the exploration campaign described above.
  • No manually labelled data was used — every training signal comes from KuaiRand-Pure's own logged interaction fields (long_view, is_click, is_like, is_follow, is_comment, is_forward, play_time_ms).

Built With

Share this project:

Updates

Submission history