Inspiration
Most recommender projects begin with a number.
A baseline gets a score. A new model gets a slightly higher score. The graph moves up, and it is easy to call that progress.
But we kept asking a harder question: what would make us trust that improvement?
A recommender can look better because a feature accidentally used future information. One random seed can look like a breakthrough. A model can be evaluated correctly and then retrained differently for submission. And when test labels exist locally, one careless code path can turn a valid experiment into an invalid result.
So we did not want to build another script that searches for a higher metric. We wanted to build a research system that could earn the right to make a claim.
That became RecommenderForge, an autonomous ML research agent for recommender systems. It turns each research idea into a traceable loop from hypothesis, to code change, to measured evidence, to a final checkpoint-backed prediction file.
What we built
RecommenderForge is an autonomous research control plane for the KuaiRand recommender benchmark.
benchmark contract
-> reproduce the official baseline
-> propose a bounded experiment
-> train and validate safely
-> verify leakage and ranking invariants
-> record metrics, failures, recovery and resources
-> compare multi-seed evidence
-> freeze the measured checkpoint
-> generate feature-only predictions
The system has four connected parts:
Benchmark adapter
Makes the dataset contract, split boundaries, evaluator, and output schema explicit. For KuaiRand-Pure, that means predicting nativelong_view, ranking only within each user’s logged impressions, and delegating GAUC and nDCG@5 calculation to the organiser’s exact evaluator.Research controller
Runs bounded modelling directions including pairwise ranking loss, leakage-safe history crosses, temporal features, and rank-space blending.Append-only experiment ledger
Records each hypothesis, parent experiment, code/configuration identity, evaluator hash, data fingerprint, metrics, resource use, failures, recovery actions, and reflection.Finalization path
Binds the selected result to the exact frozen checkpoint that earned it, then generates a feature-only prediction file without retraining or reading hidden test labels.
The contribution is not simply another recommender model. It is a system designed to make autonomous research inspectable.
Results
We started by reproducing the organiser’s five-seed NumPy Factorization Machine baseline. Only after that did the system begin ranking-oriented experiments.
| Validation result | GAUC | nDCG@5 | Primary |
|---|---|---|---|
| Organiser FM baseline | 0.667400 | 0.535744 | 0.601572 |
| RecommenderForge final blend | 0.670845 | 0.537188 | 0.604017 |
Improvement in primary: +0.002445
The selected blend combines:
BPR pairwise ranking 37.5%
history-cross features 37.5%
temporal features 25.0%
The final result was not selected from one lucky seed or one peak score. The campaign revalidated the leader and then recorded three declared non-significant confirmation batches under the same evaluator hash and data fingerprint. The selected result is tied to a frozen checkpoint, which generated a schema-valid Pure prediction file containing 170,588 feature-only test rows.
We do not claim a hidden-test score. Test labels were not used for candidate selection, local test scoring, or final-output generation. The final CSV is an inference artifact, not evidence manufactured from answers we were not meant to see.
What failed, and why we kept it
The project became more interesting when reasonable ideas failed.
A top-five lambda-weighted BPR objective reached validation primary 0.601647, below the baseline. A multi-task model using click, like, and follow as auxiliary training signals reached 0.602298, also below the selected blend.
Those experiments remain in the ledger.
A research system that remembers only successful runs is not really learning. It is editing its own history.
We also recorded a run that failed after training because of a missing NumPy import during an ordering audit. Instead of leaving behind a misleading partial result, the controller finalised it as recovered, reverted to its immutable parent, and excluded it from selection. It has no accepted metric and no checkpoint that can quietly be reused.
Failure is part of the evidence trail.
The hard engineering problems
Hidden-test integrity
The local benchmark setup makes test labels technically accessible. That does not make them valid research feedback.
Candidate-facing code fails closed if it requests test labels or tries to score the test split. Final inference receives only row identifiers and inference-known features. The final prediction file is generated from the checkpoint that was actually measured, never from a fresh retrain.
Leakage-safe history
History features can easily become invalid. A feature built from future events leaks information. A feature that is constant for a user cannot change that user’s within-user ranking at all.
RecommenderForge checks both. History features use only strictly earlier events, and every new feature family must demonstrate that it actually changes within-user ordering before it can justify another experiment.
Implementation parity
Moving from a NumPy baseline to a PyTorch ranking model could have confused an implementation change with a modelling improvement. We therefore used a pointwise PyTorch parity gate before allowing ranking-loss experiments.
Only then could a changed score be interpreted as evidence about the objective rather than an unnoticed rewrite.
Recovery and scale
Long-running ML work is where ordinary experiment scripts often break.
On KuaiRand-27K, we deliberately interrupted a 114,832,239-row feature-only output process, resumed it from a checkpoint, and verified that the recovered output was byte-identical to the uninterrupted artifact.
| Scale profile | Validation primary | Feature-only output |
|---|---|---|
| KuaiRand-1K | 0.545843 | 4,132,081 rows |
| KuaiRand-27K | 0.586756 | 114,832,239 rows |
These are not claims that we beat a bonus benchmark. The organisers have not supplied an official 1K or 27K baseline, threshold, or upload route. They are scale and recovery evidence, showing that the same safety, checkpointing, and output-integrity rules still hold when the dataset is no longer small.
Resource accounting
The confirmed campaign ran autonomously with a deterministic offline planner on a laptop:
CPU time 26.65 seconds
GPU time 0 seconds
LLM input tokens 0
LLM output tokens 0
manual interventions 0
These are real measured values, not missing measurements. The campaign did not need a GPU, and it used a deterministic planner rather than making external LLM calls. The project includes a schema-bounded provider-backed planner interface, but we do not claim it generated the final campaign decisions.
That trade-off is deliberate and visible. The system demonstrated autonomous execution, evidence preservation, recovery, convergence, and zero mid-run manual interventions. It does not pretend an LLM was used when one was not.
What we learned
The modelling lesson was clear: a pairwise ranking objective is a better fit than isolated binary prediction when the real task is to order impressions for the same user.
The bigger lesson was about reliability.
A small gain with a complete trail behind it is more valuable than a larger score that cannot be explained. Our final improvement is +0.002445, but the reason we are comfortable publishing it is the chain behind it:
same organiser evaluator
+ same declared validation protocol
+ multi-seed evidence
+ leakage controls
+ frozen checkpoint identity
+ feature-only final output
= a result another person can inspect and reproduce
The number matters. The path that produced it matters more.
What’s next
The provider-backed planner can be exercised in a bounded qualification run with the same ledger, token accounting, safety checks, and human approval gates used by the deterministic path.
We would also extend the benchmark adapter only when an official evaluator and submission route are available for another dataset. RecommenderForge should become more capable over time, but it should never become less honest while doing it.

Log in or sign up for Devpost to join the conversation.