For Agents Page - An Autonomous ML Research Agent for KuaiRand-Pure Ranking
Devpost written project description — TikTok TechJam, Track 2
How our solution addresses the problem statement
The track asks for an autonomous agent that improves a Factorisation
Machine (FM) baseline on the KuaiRand-Pure ranking task, scored by
primary = mean(GAUC, nDCG@5) on the validation split, and judged as much
on reasoning quality, autonomy, and cost-efficiency as on raw accuracy
gains.
Our answer is a three-agent loop — Research → Coding → Evaluator — wired together by a deterministic orchestrator, deliberately built so that every part of the control flow is fixed code, not an LLM's judgment call, and every part that genuinely needs reasoning gets a dedicated agent:
Research Agent (
agent/research/) proposes one concrete, literature- grounded hypothesis per iteration. It uses a breadth-then-depth flow: generate several shallow candidate directions cheaply, filter out anything matching a known dead end or a prior attempt (via a persistent, cross-run "Do/Don't" ledger and a similarity check), then write one full structured proposal for the single best survivor. Every claim is grounded in a bundled citation catalog we verified against the actual papers (not model memory), and every proposal passes a safety scan that rejects anything trying to manipulate the benchmark or leak test-split information into its own reasoning.Coding Agent (
agent/coding/) turns one hypothesis into a runnable solution directory. Before ever spending a full multi-seed training run, it runs a cheap inner loop — generate → static checks → a small smoke run → repair — so aNameErrornever costs a real attempt. Each new idea builds on the current best checkpoint rather than always starting from the static baseline, so improvements compound instead of resetting every iteration. It's permitted to use any library actually installed in the environment (measured by import availability, not a hand-maintained allowlist), so long as it stays inside a documented sandbox (no network, no subprocess spawning, one CPU core, a fixed per-run timeout).Evaluator Agent (
agent/evaluator/) runs each candidate and decides ACCEPT / REVERT / ABANDON. It never trusts a solution's self-reported score. Every result is independently re-derived from the raw persisted predictions through the competition's own unmodifiedevaluate.py(agent/verification.py), so a generatedtrain.pycan never just claim a good number. Hidden test-split metrics are quarantined at the executor level and structurally blocked from ever reaching an agent-facing payload. This ensures that no agent ever sees them, at any point.Orchestrator (
agent/orchestrator.py,agent/convergence.py) owns the parts that don't need judgment: the competition's stopping rule (50 iterations, or 6h wall-clock, or 3 consecutive iterations without a validation-primary improvement greater than 0.002 — whichever fires first), a two-tier retry/abandonment policy (3 repair attempts per idea before abandoning it; escalate to a human only after two ideas in a row fail outright), and per-iteration token/cost accounting across all three agents.
A separate one-time exploration campaign (scripts/seed_findings.py,
docs/exploration-campaign.md) runs the Starter Kit's own ranked list of
"unexplored directions" (ranking loss, sequence modeling, multi-task heads,
watch-time modeling, time drift, unbiased validation) through the real
Coding → Executor → Evaluator loop before the graded run starts, so the
Research Agent begins with an honestly measured Do/Don't ledger instead
of an empty one or a set of assumptions.
Development tools used
- Claude Code (Anthropic's CLI coding agent) — used throughout
development; visible directly in the repo's own commit history via
co-author trailers (
Claude Sonnet 5,Claude Opus 5) across most of the project. - VS Code — primary editor.
- Git / GitHub for version control and PR-based review between teammates.
- pytest as the test runner during development (717 tests as of the
pre-graded-runtag).
APIs used
- OpenAI API (
gpt-5, via theopenaiPython SDK) — the only runtime API our agent pipeline calls. All three LLM-backed agents (Research, Coding, Evaluator) share one client through dependency injection; the model is configurable viaCODING_AGENT_MODEL. An--offlinemode swaps in a deterministic, zero-cost template library and rule-based evaluator so the entire pipeline can also run end-to-end with no API key and no spend, for testing.
Libraries and frameworks used
Pinned in requirements.txt:
- NumPy (>=1.24) — the baseline FM implementation and most harness code.
- PyTorch (>=2.2) — for solutions whose hypothesis needs gradients (sequence models, multi-task heads, DeepFM-style architectures), so the Coding Agent doesn't have to hand-derive backprop.
- openai (>=1.40) — Python SDK for the GPT-5 calls above.
- PyYAML (>=6.0) — parses
solution/config.yaml; a solution directory falls back to a small stdlib YAML reader when PyYAML isn't installed, and a test asserts the two agree. - pytest (>=8.0) — the test suite.
Also available to (and usable by) generated solutions, though not
currently pinned in requirements.txt — pandas, scikit-learn, and
scipy were installed on the development machine and are measured as
available via importlib.util.find_spec rather than declared in a static
allowlist. (Flagging this as a loose end: if the graded-run environment is
provisioned strictly from requirements.txt, a solution the Coding Agent
writes expecting these libraries would fail to reproduce there — worth
pinning them explicitly before the graded run.)
Datasets and assets used
- KuaiRand-Pure (via the competition's Starter Kit, originally from Zenodo) — the sequential recommendation/ranking dataset this track is built on: ~1.14M training rows, ~125K validation rows, ~171K test rows of logged interactions, plus a separate ~1.18M-row randomized-exposure log used only as a read-only diagnostic (never as training data).
- The Starter Kit's vendored, unmodified scoring harness —
evaluate.py(the sole scoring authority for GAUC / nDCG@5 / primary),data.py(official date-based train/valid/test splits and feature encoding), andbaseline.py(the FM/pop/random baselines we measure against) — copied byte-for-byte with SHA-256 provenance tracking so we can prove none of it was modified. - A bundled literature citation catalog (
agent/research/references.json) grounding the Research Agent's proposals — every entry checked against the actual paper text (e.g., Covington et al. 2016 on YouTube recommendations; Schnabel et al. 2016 on propensity-weighted evaluation) rather than trusted from model memory. - A persistent cross-run findings ledger
(
agent/research/findings.jsonl), seeded by the exploration campaign described above. - No manually labelled data was used — every training signal comes
from KuaiRand-Pure's own logged interaction fields (
long_view,is_click,is_like,is_follow,is_comment,is_forward,play_time_ms).
Log in or sign up for Devpost to join the conversation.