REX: Evidence-Audited Autonomous Recommender Research Agent

Inspiration

Recommendation research is a repeated scientific loop: understand the data, form a hypothesis, change a model or feature, train it, evaluate the result, diagnose what happened, and decide what to try next. An LLM can help automate that loop, but an uncontrolled coding agent can leak validation information, modify the evaluator, lose a good checkpoint, or waste hours after a crash.

We built REX to make autonomy and scientific validity part of the same system. The LLM acts as the researcher, while trusted code controls the data boundary, evaluation, experiment state, budgets, recovery, and final submission.

What it does

REX autonomously:

  1. verifies the official KuaiRand-Pure files, temporal split, evaluator, and five-seed baseline;
  2. selects a recommendation-research hypothesis and asks an authorized LLM for a structured proposal or constrained patch;
  3. rejects protected-file changes and unsafe capabilities before training;
  4. runs a small temporal-shadow experiment before spending resources on three complete temporal folds;
  5. compares each candidate with a matched control and records GAUC, nDCG@5, uncertainty, diagnostics, and artifacts;
  6. automatically recovers durable work after a controller or worker failure;
  7. preserves the validation-best model rather than simply keeping the last model attempted; and
  8. stops through convergence, the 50-iteration cap, or the six-hour ceiling.

Test prediction is a separate, explicitly authorized operation. The finalizer loads only the immutable winning model, predicts the canonical 170,588 test rows, rejects invalid or misaligned output, runs the organizer's checker twice, and seals the model, predictions, reports, configuration, source identity, and hashes. Hidden test labels are never loaded or scored locally.

How we built it

REX uses a trusted-controller/untrusted-worker architecture. The controller owns the experiment database, data permissions, evaluator, convergence rule, and evidence ledger. Generated candidate code runs in disposable Docker workers that are non-root, networkless, credential-free, resource-bounded, and given only explicit filesystem access.

Durable state is stored in SQLite with a hash-chained event history. Every experiment records its hypothesis, parent, effective configuration, code or configuration change, metrics, diagnostics, failures, repair attempts, checkpoint, Git commit, dependency identity, and Docker image digest. Cheap tests reject weak ideas early, while promising candidates advance to three temporal shadow folds and official validation.

The production researcher used a locally authenticated OpenAI Codex CLI. REX also supports Claude CLI, explicitly authorized OpenAI API access, and a fixed experiment queue that requires no LLM. The LLM does not make video predictions; it manages the research process. NumPy/Factorization Machine and LightGBM model plugins perform the actual learning.

Autonomous run and recovery

The selected immutable run was r3-docker-20260831-codex-v15. During its first iteration, the rehearsal deliberately terminated an active controller after a worker lease had been stored durably. REX detected the stale owner, recovered the leased work, resumed under the original deadline, preserved the incumbent and experiment counters, and completed with zero manual interventions.

The run evaluated three scored iterations and stopped using the preconfigured epsilon=0.002, N=3 cumulative plateau rule. The winning E15 experiment isolated one change: aggregate five context-aware Factorization Machine members by arithmetic mean instead of median. The members use seeds 0–4, latent dimension 16, learning rate 0.001, L2 regularization of 0.000001, seven epochs, and batch size 8192.

Results

The primary score is the arithmetic mean of GAUC and nDCG@5.

KuaiRand-Pure validation result GAUC nDCG@5 Primary
Organizer official FM baseline 0.667400 0.535700 0.601600
REX validation-best model 0.670204 0.537110 0.603657
Absolute improvement +0.002804 +0.001410 +0.002057

These are validation metrics. We do not claim a hidden-test score. The final CSV contains exactly 170,588 rows with row_id,user_id,video_id,score, and the organizer's format/alignment checker passed twice with test_scored: false.

Resource usage and autonomy

Resource Recorded usage
Agent wall-clock to convergence 1,484.200829 seconds (24 min 44.2 s)
Iterations 3 / 50
LLM calls 11
LLM input tokens 289,862
LLM output tokens 20,949
Total LLM tokens 310,811
GPU-hours 0.0
Manual interventions 0

What makes it different

The main contribution is not one model class. It is an inspectable research system that combines matched treatment/control experiments, temporal shadow folds, constrained agent patches, immutable model bundles, automatic recovery, protected evaluation, test-label isolation, resource accounting, and a sealed one-time submission handoff. The evidence shows not just the final number, but what the agent tried, why it tried it, what changed, what failed, and how the system recovered.

Challenges we encountered

The hardest work was creating a trustworthy boundary around autonomous code. We had to make process interruption, stale leases, timeouts, invalid predictions, concurrent workers, path translation between macOS and Docker, and final artifact copying recoverable without allowing a failed candidate to replace the current champion. We also had to keep temporal feature computation strictly separated from validation and test labels.

Limitations and future improvements

  • The final run converged after three scored iterations and did not exhaust the wider feature/model queue.
  • The winning model contains five deterministic ensemble members, but the full outer run was not repeated under several independent run seeds.
  • Only KuaiRand-Pure was submitted; the two optional bonus benchmarks were not attempted.
  • Offline GAUC and nDCG@5 do not guarantee the same gain in an online system.

With more time, we would extend the controlled search across inference-safe metadata, temporal history, multi-feedback objectives, sequence models, and censored watch-time losses; confirm finalists across additional seeds; and run the sealed workflow on KuaiRand-1k and KuaiRand-27k.

Tools, APIs, libraries, and data

  • Development and operations: Git, Python 3.13, Docker Desktop/Buildx, SQLite, VS Code/Codex desktop, pytest, and Ruff.
  • LLM researcher used in the final run: locally authenticated OpenAI Codex CLI with strict structured outputs. Claude CLI and explicitly authorized OpenAI API access are also supported.
  • ML and data: NumPy, pandas, LightGBM, scikit-learn, PyYAML, and the frozen organizer evaluation/submission scripts.
  • Dataset: only the official KuaiRand-Pure train and validation data for model development. KuaiRand-1k, KuaiRand-27k, randomized-exposure logs, external training data, and hidden-test labels were not used for this result.

Team contributions

  • Thangaraju Sibiraj (also recorded as Sibi in earlier Git commits): led the production implementation, Docker isolation, scientific experiment and model/feature systems, fault recovery, evidence reporting, documentation, and final submission workflow.
  • Amudhan: co-designed the project and experiment strategy, contributed to the system and modeling discussions, and built the initial autonomous harness, including convergence and budget logic, the first training/orchestration path, schemas, and submission validation.
  • Together: we shaped the research direction, discussed candidate ideas and tradeoffs, reviewed the system's behaviour, and interpreted the experimental evidence.

Built With

Share this project:

Updates

Submission history