REX: Evidence-Audited Autonomous Recommender Research Agent
Inspiration
Recommendation research is a repeated scientific loop: understand the data, form a hypothesis, change a model or feature, train it, evaluate the result, diagnose what happened, and decide what to try next. An LLM can help automate that loop, but an uncontrolled coding agent can leak validation information, modify the evaluator, lose a good checkpoint, or waste hours after a crash.
We built REX to make autonomy and scientific validity part of the same system. The LLM acts as the researcher, while trusted code controls the data boundary, evaluation, experiment state, budgets, recovery, and final submission.
What it does
REX autonomously:
- verifies the official KuaiRand-Pure files, temporal split, evaluator, and five-seed baseline;
- selects a recommendation-research hypothesis and asks an authorized LLM for a structured proposal or constrained patch;
- rejects protected-file changes and unsafe capabilities before training;
- runs a small temporal-shadow experiment before spending resources on three complete temporal folds;
- compares each candidate with a matched control and records GAUC, nDCG@5, uncertainty, diagnostics, and artifacts;
- automatically recovers durable work after a controller or worker failure;
- preserves the validation-best model rather than simply keeping the last model attempted; and
- stops through convergence, the 50-iteration cap, or the six-hour ceiling.
Test prediction is a separate, explicitly authorized operation. The finalizer loads only the immutable winning model, predicts the canonical 170,588 test rows, rejects invalid or misaligned output, runs the organizer's checker twice, and seals the model, predictions, reports, configuration, source identity, and hashes. Hidden test labels are never loaded or scored locally.
How we built it
REX uses a trusted-controller/untrusted-worker architecture. The controller owns the experiment database, data permissions, evaluator, convergence rule, and evidence ledger. Generated candidate code runs in disposable Docker workers that are non-root, networkless, credential-free, resource-bounded, and given only explicit filesystem access.
Durable state is stored in SQLite with a hash-chained event history. Every experiment records its hypothesis, parent, effective configuration, code or configuration change, metrics, diagnostics, failures, repair attempts, checkpoint, Git commit, dependency identity, and Docker image digest. Cheap tests reject weak ideas early, while promising candidates advance to three temporal shadow folds and official validation.
The production researcher used a locally authenticated OpenAI Codex CLI. REX also supports Claude CLI, explicitly authorized OpenAI API access, and a fixed experiment queue that requires no LLM. The LLM does not make video predictions; it manages the research process. NumPy/Factorization Machine and LightGBM model plugins perform the actual learning.
Autonomous run and recovery
The selected immutable run was r3-docker-20260831-codex-v15. During its first
iteration, the rehearsal deliberately terminated an active controller after a
worker lease had been stored durably. REX detected the stale owner, recovered
the leased work, resumed under the original deadline, preserved the incumbent
and experiment counters, and completed with zero manual interventions.
The run evaluated three scored iterations and stopped using the preconfigured
epsilon=0.002, N=3 cumulative plateau rule. The winning E15 experiment
isolated one change: aggregate five context-aware Factorization Machine members
by arithmetic mean instead of median. The members use seeds 0–4, latent
dimension 16, learning rate 0.001, L2 regularization of 0.000001, seven epochs,
and batch size 8192.
Results
The primary score is the arithmetic mean of GAUC and nDCG@5.
| KuaiRand-Pure validation result | GAUC | nDCG@5 | Primary |
|---|---|---|---|
| Organizer official FM baseline | 0.667400 | 0.535700 | 0.601600 |
| REX validation-best model | 0.670204 | 0.537110 | 0.603657 |
| Absolute improvement | +0.002804 | +0.001410 | +0.002057 |
These are validation metrics. We do not claim a hidden-test score. The final
CSV contains exactly 170,588 rows with
row_id,user_id,video_id,score, and the organizer's format/alignment checker
passed twice with test_scored: false.
Resource usage and autonomy
| Resource | Recorded usage |
|---|---|
| Agent wall-clock to convergence | 1,484.200829 seconds (24 min 44.2 s) |
| Iterations | 3 / 50 |
| LLM calls | 11 |
| LLM input tokens | 289,862 |
| LLM output tokens | 20,949 |
| Total LLM tokens | 310,811 |
| GPU-hours | 0.0 |
| Manual interventions | 0 |
What makes it different
The main contribution is not one model class. It is an inspectable research system that combines matched treatment/control experiments, temporal shadow folds, constrained agent patches, immutable model bundles, automatic recovery, protected evaluation, test-label isolation, resource accounting, and a sealed one-time submission handoff. The evidence shows not just the final number, but what the agent tried, why it tried it, what changed, what failed, and how the system recovered.
Challenges we encountered
The hardest work was creating a trustworthy boundary around autonomous code. We had to make process interruption, stale leases, timeouts, invalid predictions, concurrent workers, path translation between macOS and Docker, and final artifact copying recoverable without allowing a failed candidate to replace the current champion. We also had to keep temporal feature computation strictly separated from validation and test labels.
Limitations and future improvements
- The final run converged after three scored iterations and did not exhaust the wider feature/model queue.
- The winning model contains five deterministic ensemble members, but the full outer run was not repeated under several independent run seeds.
- Only KuaiRand-Pure was submitted; the two optional bonus benchmarks were not attempted.
- Offline GAUC and nDCG@5 do not guarantee the same gain in an online system.
With more time, we would extend the controlled search across inference-safe metadata, temporal history, multi-feedback objectives, sequence models, and censored watch-time losses; confirm finalists across additional seeds; and run the sealed workflow on KuaiRand-1k and KuaiRand-27k.
Tools, APIs, libraries, and data
- Development and operations: Git, Python 3.13, Docker Desktop/Buildx, SQLite, VS Code/Codex desktop, pytest, and Ruff.
- LLM researcher used in the final run: locally authenticated OpenAI Codex CLI with strict structured outputs. Claude CLI and explicitly authorized OpenAI API access are also supported.
- ML and data: NumPy, pandas, LightGBM, scikit-learn, PyYAML, and the frozen organizer evaluation/submission scripts.
- Dataset: only the official KuaiRand-Pure train and validation data for model development. KuaiRand-1k, KuaiRand-27k, randomized-exposure logs, external training data, and hidden-test labels were not used for this result.
Team contributions
- Thangaraju Sibiraj (also recorded as Sibi in earlier Git commits): led the production implementation, Docker isolation, scientific experiment and model/feature systems, fault recovery, evidence reporting, documentation, and final submission workflow.
- Amudhan: co-designed the project and experiment strategy, contributed to the system and modeling discussions, and built the initial autonomous harness, including convergence and budget logic, the first training/orchestration path, schemas, and submission validation.
- Together: we shaped the research direction, discussed candidate ideas and tradeoffs, reviewed the system's behaviour, and interpreted the experimental evidence.
Log in or sign up for Devpost to join the conversation.