中文版本请见文末 · Chinese version below

About the project

Machine-learning engineers spend a surprising amount of time doing the same loop over and over again: read the problem, inspect the data, change the model or features, train, evaluate, reflect, and try again. The TikTok TechJam recommender-systems challenge made that loop especially clear because the task is not just to train a single model, but to improve over an official baseline on a real recommendation benchmark under clear constraints. That inspired RecPilot: an autonomous ML research agent for recommender systems that can run this experimentation loop with minimal human intervention.

The core idea behind RecPilot is simple: instead of manually editing code for every experiment, an agent proposes the next change, applies it to a safe experiment configuration, launches training, evaluates the result, decides whether to keep or roll back the change, logs the outcome, and repeats. In practice, that means RecPilot behaves like a small, tireless research assistant for ranking systems. It does not replace judgment entirely, but it does automate a large part of the repetitive engineering work that slows recommender-system iteration.

What the project does

RecPilot is designed for the KuaiRand-Pure benchmark in the TikTok TechJam Track 2 problem. The system starts from the official Factorization Machine baseline and then iteratively explores stronger ranking pipelines. It can modify feature engineering, loss functions, model families, and hyperparameters while keeping the evaluation protocol fixed and leakage-safe.

At a high level, the system runs this loop:

  1. Read the current state of experiments.
  2. Propose the next experiment hypothesis.
  3. Apply an allowed operator such as adding history features, recency features, or switching the ranking objective.
  4. Train and evaluate the new configuration.
  5. Keep the new configuration if it improves validation primary score; otherwise roll back.
  6. Log the metrics, decision, runtime, and any recovery actions.
  7. Repeat until convergence or budget limits are reached.

The optimization target is the challenge primary metric:

\( \text{Primary Score} = \frac{\text{GAUC} + \text{nDCG@5}}{2} \)

This is important because the agent is not just trying to optimize a training loss; it is trying to make decisions that improve the actual competition objective.

How it was built

RecPilot is built as a modular experimentation system rather than a monolithic script.

1. Planner

The planner chooses the next experiment. In one mode, an LLM proposes the next operator and its parameters based on the current best run, recent outcomes, and what has already been tried. In fallback mode, a heuristic planner follows a structured priority order so the system still works without an API key.

2. Operator catalog

Instead of letting the agent write arbitrary code, RecPilot gives it a catalog of safe, meaningful changes. These include:

  • Reproducing the official FM baseline
  • Adding user-history crosses
  • Adding recency-weighted history features
  • Trying sequence-interest modeling
  • Switching ranking losses
  • Tuning hyperparameters
  • Blending with popularity-based signals

This makes the system much more controllable and auditable.

3. Harness

The harness is the execution engine for a single experiment. Given a config, it:

  • Loads the benchmark data
  • Applies feature transforms
  • Encodes categorical fields
  • Trains the selected model
  • Evaluates it with the official metrics
  • Writes a result.json and submission file

This separation turned out to be crucial. The planner decides what to try, and the harness decides how to run it reproducibly.

4. Autonomous loop

The loop ties everything together. It repeatedly asks for the next experiment, launches it in a subprocess, reads the result, compares it to the current best validation score, keeps or rolls back the change, and records an event in the run log. It also handles timeouts, retries, cooldowns for failing operators, and stop conditions such as iteration caps or convergence.

5. Reporting and audit

Every iteration is logged with its hypothesis, config change, metrics, and decision. That means the project is not just a final score; it is an auditable research trail. Multi-seed evaluation support was also added so promising configurations can be checked for stability rather than relying on one lucky run.

What was learned

This project taught several useful lessons about recommender systems and about autonomous research workflows.

First, good recommender improvements are often not flashy architecture changes. Small, targeted feature ideas such as user-history crosses and recency-weighted interactions can matter a lot because recommendation is deeply contextual and time-sensitive.

Second, autonomy is only useful if it is constrained well. An unconstrained agent that edits arbitrary code can become brittle very quickly. A constrained agent with a strong operator catalog, reproducible harness, rollback behavior, and logs is much more practical.

Third, evaluation discipline matters as much as modeling. For this challenge, it was critical to preserve the official split, avoid temporal leakage, and use validation only for model selection. That influenced the entire system design.

Challenges faced

The hardest challenge was not writing one model. It was building a reliable experimental system.

One challenge was deciding how much freedom to give the agent. Too little freedom makes the system feel scripted; too much freedom makes it unstable. The solution was to define a structured operator catalog so the agent could still make meaningful choices without breaking the pipeline.

Another challenge was handling failure during long runs. Training jobs can timeout, produce weak results, or fail unexpectedly. To address this, the loop was designed with rollback, retries, cooldowns, subprocess isolation, and persistent run state so a long experiment session would not collapse after one bad iteration.

A third challenge was balancing ambition with realism. It is easy to keep adding more model families, but under hackathon time pressure the more important problem is turning experimentation into something that is robust, reproducible, and explainable.

Current result and project direction

RecPilot already reproduces the official baseline and has begun improving on it by introducing leakage-safe history and recency features. More importantly, it turns the competition workflow itself into a system: hypothesis generation, experiment execution, evaluation, recovery, and reporting all happen inside one reproducible loop.

That is the bigger idea behind the project. The goal is not only to submit a stronger recommender model, but to show what a practical autonomous ML research workflow can look like for recommender-system development.

Why this matters

Recommendation teams often spend huge amounts of time on iterative model engineering. If that loop can be partially automated in a robust and auditable way, the benefit is bigger than one benchmark score. It means faster iteration, clearer experiment history, lower manual overhead, and a more realistic path toward AI systems that assist real ML engineers instead of just generating isolated code snippets.

RecPilot is an attempt to make that future concrete.


中文版本 / Chinese Version

关于项目

机器学习工程师有大量时间花在同一个循环上:阅读问题、检查数据、修改模型或特征、训练、评估、复盘、再试一次。TikTok TechJam 推荐系统赛题让这个循环格外清晰,因为任务不只是训练单个模型,而是在明确约束下,在真实推荐基准上超越官方基线。这启发了 RecPilot:一个面向推荐系统的自主机器学习研究智能体,能够在极少人工干预的情况下运行这个实验循环。

RecPilot 的核心想法很简单:不再为每次实验手动改代码,而是由智能体提出下一步改动、将其应用到安全的实验配置、启动训练、评估结果、决定保留还是回滚、记录结果,然后重复。实际效果是,RecPilot 像一个小而不知疲倦的排序系统研究助手。它不会完全取代判断力,但确实自动化了拖慢推荐系统迭代的大部分重复性工程工作。

项目功能

RecPilot 针对 TikTok TechJam Track 2 赛题中的 KuaiRand-Pure 基准设计。系统从官方的因子分解机(FM) 基线出发,迭代探索更强的排序流水线。它可以修改特征工程、损失函数、模型族和超参数,同时保持评估协议固定且无泄露。

系统的高层循环如下:

  • 读取当前实验状态
  • 提出下一个实验假设
  • 应用一个允许的算子,例如加入历史特征、时近性特征,或切换排序目标
  • 训练并评估新配置
  • 若验证集主指标提升则保留新配置,否则回滚
  • 记录指标、决策、运行时长以及任何恢复动作
  • 重复,直至收敛或达到预算上限

优化目标是赛题的主指标。这一点很重要,因为智能体不只是在优化训练损失,而是在做能真正提升比赛目标的决策。

构建方式

RecPilot 被构建为一个模块化的实验系统,而不是单体脚本。

1. 规划器

规划器决定下一个实验。在一种模式下,由 LLM 根据当前最优运行、近期结果以及已尝试过的内容,提出下一个算子及其参数。在回退模式下,启发式规划器按结构化的优先级顺序执行,因此没有 API key 时系统依然可用。

2. 算子目录

RecPilot 不让智能体编写任意代码,而是提供一份安全且有意义的改动目录,包括:

  • 复现官方 FM 基线
  • 加入用户历史交叉特征
  • 加入时近性加权的历史特征
  • 尝试序列兴趣建模
  • 切换排序损失
  • 调整超参数
  • 与热度类信号做融合

这让系统更可控、更可审计。

3. 实验执行器

实验执行器是单次实验的执行引擎。给定一份配置,它会:

  • 加载基准数据
  • 应用特征变换
  • 编码类别字段
  • 训练所选模型
  • 使用官方指标进行评估
  • 写出 result.json 与提交文件

这个分离非常关键:规划器决定试什么,执行器决定如何可复现地运行。

4. 自主实验循环

循环把所有部分串起来。它反复请求下一个实验、在子进程中启动、读取结果、与当前最优验证分数比较、保留或回滚改动,并在运行日志中记录一条事件。它同时处理超时、重试、失败算子的冷却,以及迭代上限或收敛等停止条件。

5. 报告与审计

每次迭代都会记录其假设、配置改动、指标和决策。因此这个项目不只是一个最终分数,而是一条可审计的研究轨迹。系统还加入了多随机种子评估支持,可以检验有潜力的配置是否稳定,而不是依赖一次幸运的运行。

学到了什么

这个项目在推荐系统和自主研究流程两方面都带来了有用的经验。

第一,好的推荐系统改进往往不是花哨的架构改动。像用户历史交叉特征、时近性加权交互这类小而精准的特征想法可能影响很大,因为推荐高度依赖上下文且对时间敏感。

第二,自主性只有在被良好约束时才有用。一个可以任意改代码的无约束智能体会很快变得脆弱。相比之下,一个带有完善算子目录、可复现执行器、回滚机制和日志的受约束智能体要实用得多。

第三,评估纪律与建模同等重要。在这个赛题中,保持官方划分、避免时间泄露、只用验证集做模型选择至关重要。这一点影响了整个系统设计。

遇到的挑战

最难的挑战不是写出一个模型,而是构建一个可靠的实验系统。

一个挑战是决定给智能体多大自由度。自由度太小,系统显得像写死的脚本;太大,则会变得不稳定。解决方案是定义结构化的算子目录,让智能体仍能做出有意义的选择,同时不会破坏流水线。

另一个挑战是处理长时间运行中的失败。训练任务可能超时、产出弱结果或意外失败。为此,循环在设计上带有回滚、重试、冷却、子进程隔离和持久化运行状态,使得长时间的实验会话不会因为一次糟糕的迭代而崩溃。

第三个挑战是在野心与现实之间取得平衡。不断加入更多模型族很容易,但在黑客松的时间压力下,更重要的问题是把实验过程变成鲁棒、可复现且可解释的东西。

当前结果与项目方向

RecPilot 已经复现了官方基线,并通过引入无泄露的历史与时近性特征开始超越它。更重要的是,它把比赛工作流本身变成了一个系统:假设生成、实验执行、评估、恢复与报告全部发生在一个可复现的循环中。

这才是项目背后更大的想法。目标不只是提交一个更强的推荐模型,而是展示面向推荐系统开发的实用自主机器学习研究工作流可以是什么样子。

为什么这很重要

推荐团队常常在迭代式模型工程上花费大量时间。如果这个循环能够以鲁棒且可审计的方式被部分自动化,收益就不止于一个基准分数。它意味着更快的迭代、更清晰的实验历史、更低的人工开销,以及一条更现实的路径——通向真正辅助机器学习工程师的 AI 系统,而不是只会生成孤立代码片段的工具。

RecPilot 正是把这个未来变得具体的一次尝试。

Built With

  • agentic-ai
  • autonomous-agents
  • factorization-machines
  • feature-engineering
  • github
  • llm
  • numpy
  • openai-api
  • python
Share this project:

Updates

Submission history