Inspiration
ML research is often reduced to a search over predefined hyperparameters. That works when the important decisions are already known, but it cannot invent a new feature, change the training objective, or decide that a different part of the ranking problem deserves more attention.
Kopibara RLab Agent treats ML engineering as a search over executable code. The agent starts with a runnable candidate, proposes one focused hypothesis, changes the code, measures the result, and uses that evidence to choose its next experiment.
The goal is to give an AI system a bounded research loop with an explicit record of what it tried, what worked, what failed, and when further search stopped producing meaningful progress.
What it does
Kopibara runs an autonomous MLE iteration loop on the KuaiRand recommendation benchmarks.
Each iteration follows the same process:
Read benchmark contract
↓
Choose a parent candidate
↓
Propose one testable hypothesis
↓
Apply one exact code patch
↓
Run candidate in a sandboxed subprocess
↓
Measure validation metrics
↓
Keep the strongest candidate
↓
Continue until convergence
The agent receives the benchmark definition, published baselines, the current solution tree, and the results of previous experiments. It proposes one change to history_lgbm.py. The patch must match its anchor text exactly once, which keeps the change precise and reviewable.
The candidate is compiled and executed with a time limit. The runner reads GAUC and nDCG@5 from validation output, records the unified diff and hypothesis, and adds the experiment to the search tree. The highest-scoring measured node becomes the parent for the next iteration.
A failed candidate does not disappear. The system records the failure, sends the original hypothesis and error text through one repair attempt, and marks the branch as recovered_failure if the repair also fails. Failed branches are excluded from future parent selection.
The agent stops when the cumulative best score remains within the convergence threshold for the configured number of iterations. It then writes the winning source, model artifacts, a manifest, and a validated submission.csv.
KuaiRand-Pure result
| Metric | Official FM baseline | Kopibara RLab Agent | Absolute delta |
|---|---|---|---|
| GAUC | 0.6674 | 0.7059 | +0.0385 |
| nDCG@5 | 0.5357 | 0.5538 | +0.0181 |
| Primary score | 0.6016 | 0.6299 | +0.0283 |
The primary score is the equal-weighted mean of GAUC and nDCG@5.
The converged Pure run used three search iterations, 5 minutes and 5 seconds of wall-clock time, 12,901 language-model tokens, zero GPU hours, and zero manual interventions.
KuaiRand-1K result
The same pipeline was retrained and evaluated on the KuaiRand-1K splits:
| Metric | Kopibara RLab Agent |
|---|---|
| GAUC | 0.7107 |
| nDCG@5 | 0.7888 |
| Primary score | 0.7498 |
The starter kit publishes no official KuaiRand-1K baseline, so no official improvement claim is made. The measured FM runtime reference was GAUC 0.6662, nDCG@5 0.5496, and primary 0.6079. That reference is not an organizer-published 1K baseline.
How we built it
The system is written in Python and organized around a small set of focused modules:
research_agent.pycontrols best-first search, recovery, convergence checks, logging, and the final manifest.planner.pybuilds prompts, parses plans, validates hypotheses, and applies exact-match code edits.runner.pycompiles candidates, runs them in bounded subprocesses, parses metrics, and creates the final submission.models.pydefines the value objects used for plans, nodes, metrics, and run records.constants.pystores model settings, budgets, convergence rules, and the token denylist.experiments/history_lgbm.pycontains the seed candidate edited by the autonomous controller.
The seed candidate is a LightGBM ranker using lagged user, item, author, and user-item feedback aggregates. It is intentionally ordinary. That gives the agent a measurable starting point and makes it possible to separate the agent's contribution from a highly specialized initial model.
The action space is code. A candidate can change the objective, ranking behavior, feature construction, or training schedule through a controlled source patch.
The planning prompt includes the full solution tree, including previous hypotheses, patches, scores, and failed branches. This gives the agent a record of what has already been tested and prevents it from repeatedly exploring the same direction.
Generated edits must match their anchor text exactly once. Each accepted iteration produces a real unified diff, so the run log captures both the hypothesis and the precise source change used to test it.
We added boundaries around generated code. The planner can edit only history_lgbm.py. Candidates cannot introduce tokens such as subprocess, socket, requests, urllib, httpx, openai, os.system, eval(, or exec( when those capabilities were not already present in the parent candidate. The API key is removed before candidate execution.
Every iteration produces a JSON record containing the hypothesis, source diff, validation metrics, recovery events, token counts, and execution state.
Challenges we ran into
The first challenge was giving the agent enough freedom to discover meaningful ML changes without allowing candidate code to become unbounded or difficult to inspect. Restricting edits to one runnable script made the search safer and easier to review, but also limited the parts of the pipeline the agent could explore.
Failure recovery was another important problem. Generated code can fail because of syntax errors, invalid assumptions about the data, timeouts, missing metric output, or runtime exceptions. The controller needed to recover where possible without hiding failures or repeatedly expanding a broken branch.
Convergence required care as well. A run can stop too early on a flat validation landscape, or select a lucky score after too many trials. Kopibara uses the organizer's rule of ε = 0.002 over three consecutive iterations and reports the converged result rather than selecting the highest score from an unlimited search.
Metric alignment was the most interesting modelling challenge. The benchmark uses both GAUC and nDCG@5, so a candidate can improve one metric while hurting the other. The agent had to reason about the evaluator's scoring behavior rather than optimize an unrelated training statistic.
Data leakage was treated as an engineering challenge rather than an assumption. Test labels and test feedback are stripped before generated code receives the data, and test predictions are never used for model selection.
Accomplishments that we're proud of
The strongest result is the improvement over the published KuaiRand-Pure FM baseline. Kopibara raised the primary validation score from 0.6016 to 0.6299, an absolute gain of 0.0283.
The winning candidate came from a metric-specific hypothesis. Because the evaluator scores nDCG@5, the agent tested whether truncating LambdaRank at five positions would concentrate the training signal on the part of the ranking that matters most.
The latest converged search trace was:
| Iteration | Hypothesis | Validation primary | Kept |
|---|---|---|---|
| 0 | Seed candidate with lagged multi-feedback history features | 0.6299 | ✓ |
| 1 | Replace LambdaRank with the XE-NDCG ranking objective | 0.6299 | ✗ |
| 2 | Widen truncation to 10 to trade nDCG@5 for GAUC | 0.6264 | ✗ |
| 3 | Disable per-query lambda normalization to match impression-weighted GAUC | 0.6268 | ✗ |
The result did not come from a large hyperparameter sweep. It came from testing a focused hypothesis, probing nearby alternatives, and stopping when the measured evidence supported convergence.
We are also proud of the operational guarantees:
- Zero manual interventions in the reported Pure run.
- No hidden test labels or test feedback passed to generated code.
- No test metrics used for model selection.
- Candidate execution bounded by subprocess timeouts.
- API credentials removed before candidate execution.
- Failed candidates repaired once and retained in the run history.
- Every experiment captured as a reviewable diff with its measured result.
- Final output includes a row-aligned, schema-checked submission.
- The same pipeline also produced a 0.7498 primary validation score on KuaiRand-1K.
The project produces both a model result and an auditable research trace. The final run contains the hypotheses, source changes, scores, recovery events, token usage, winning candidate, and submission artifacts.
What we learned
The biggest lesson was that objective alignment mattered more than additional model complexity in this setting. The winning Pure result did not require a larger architecture or an expensive training setup. It changed how the existing ranker spent its optimization signal.
We also learned that code search gives the agent a larger action space than hyperparameter search. A single source change can modify features, loss functions, ranking cutoffs, or training behavior while keeping the evaluation loop unchanged.
A strong seed candidate matters. If the starting point is too weak, the agent may spend its budget fixing basic implementation problems. If it is too specialized, improvements become difficult to attribute. The lagged-history LightGBM candidate provided a practical middle ground.
Failure handling is part of autonomous research. A system that records only successful candidates hides important evidence. Keeping failed branches allows the planner to avoid repeating them and gives the final run a more honest account of the search.
We also learned that convergence is sensitive to the shape of the validation landscape. The Pure run reached its best result at the seed candidate, while later hypotheses helped confirm that nearby objective changes did not improve it. This makes the stopping rule useful, but also highlights the risk of ending before a very different branch has been explored.
The KuaiRand-1K result shows that the pipeline can transfer beyond the required Pure benchmark, while the lack of an official 1K baseline means that result must be presented as an absolute validation score rather than a claimed improvement.
What's next for Kopibara RLab Agent
The current controller edits one candidate file. The next step is to expand the edit boundary to the feature-construction module and other parts of the training pipeline. This would increase the search space while requiring stronger patch validation and a larger token denylist.
We also want to support parallel candidate execution. Candidates are independent subprocesses, so several branches could be explored at once within the same wall-clock budget.
The convergence logic could gain a minimum iteration floor before the convergence window becomes active. This would reduce the chance that a flat run stops before a different model family produces a useful result.
KuaiRand includes randomized-exposure logs that could support off-policy evaluation and debiasing experiments. These logs are outside the current training path, but they could inform a future validation-time sanity check or exposure-aware objective.
Deep models such as deepfm_mtl.py and din_ranker.py are already available, but neither beat the gradient-boosted ranker on Pure at the current scale. A larger search budget and better allocation of training time would provide a stronger comparison.
The long-term direction is an autonomous ML research system that can explore more of the pipeline while preserving the properties demonstrated here: focused hypotheses, executable experiments, measurable evidence, recoverable failures, and a complete record of how the final result was reached.


Log in or sign up for Devpost to join the conversation.