Devpost — paste-ready (standard story fields)

Copy each heading into the matching Devpost box. Metrics are kit GAUC / nDCG@5 (primary = mean). Not NDCG@10 / Recall@50.

Title: recagent — Autonomous HPO research agent for KuaiRand-Pure within-user ranking

Tagline: Autonomous FM HPO for KuaiRand-Pure: Qwen LoRA + GP-EI (HPUCB) search lr/l2/k on date-split valid; ship valid-best test ranks. Modest gain over official FM.


Inspiration

I wanted an agent that behaves like a small research loop—propose, train, score, reflect—without a coordinator LLM, without leaking test into selection, and and instead of using pure bayesian optimization algorithms I tested and proved that mixing them with trained LLMs can do a better job at it.


What it does

recagent runs:

Read → EDA → Retrieve → FE → CASH → Loss → HPO → Eval → OPE → Reflect → loop

This submission is Piece 1: FE, CASH, and loss stay frozen (official 5 fields + FM + logloss). The agent searches FM hyperparameters (lr, l2, k, epochs, patience, batch_size).

  • Train dates (20220408–20220421): fit FM and vocab only.
  • Valid/eval dates (20220422–20220428): HPO score and incumbent. Full official valid (22,377 users / 124,909 rows).
  • Test dates (20220429–20220508): scored once at the valid-best checkpoint (23,875 users / 170,588 rows). Never used to pick HPs.
  • Splits are by date. The same user can appear in train and valid on different days.

HPO: HPUCB_arm chooses a Qwen3.5-9B LoRA arm vs GP-EI (not TPE). Iter 0 is official FM. Each trial is logged (iterations.jsonl). Finish writes a kit CSV (row_id,user_id,video_id,score). OPE on log_random_* ∩ valid is logged only, not used for selection.

Submitted run: ORPO LoRA + HPUCB, 80 trials, seed 1, full Pure. Test primary 0.5969 (GAUC 0.6636, nDCG@5 0.5303) vs official pin 0.5946. Kit --check passed.

Method Valid Test @ valid-best Δ test
Official FM 0.6015 0.5946 —
80-trial GP-EI only 0.6031 0.5968 +0.0022
SFT LoRA + HPUCB 0.6030 0.5960 +0.0014
ORPO LoRA + HPUCB (submit) 0.6032 0.5969 +0.0023

The incumbent is GP-EI (iter 66), not the LoRA proposal. The ORPO arm still beat official on valid (~0.6029). Hybrid is ~0.0001 above matched 80-BO at n=1.


How we built it

Python orchestrator + numpy FM (CPU). No pandas/sklearn/torch on the default agent path. Kit evaluate.py / submit.py for metrics and CSV.

Loop (Piece 1): Retrieve (RS3f05, trials 1–3). FE/CASH/loss frozen. HPUCB ω_llm 0.9→0.1. P5+ history (top/worst/recent). Reflect = gpt-4o-mini via an OpenAI-compatible gateway. HPO proposals = local Qwen3.5-9B + LoRA via LlamaFactory (HPO_BASE_URL). --no-llm maps the LLM arm to GP-EI.

LoRA (proposer, not ranking labels): SFT on HPO prompt → next-config from the experience bank (valid primary; test never in LoRA). Then ORPO on the SFT adapter.

  • SFT: r=32, α=64, dropout=0.05, LoRA+ 16, NEFTune α=5, lr=1e-4, 3 epochs, cutoff 2048, ~2,775 examples, 261 steps, 2× A800.
  • ORPO: β=0.1, lr=1e-5, 1 epoch, 705 pairs, 45 steps.

Data: KuaiRand-Pure + an internal BO experience bank. ~13.5 GPU-h for SFT+ORPO+serve; FM HPO is CPU.

Artifacts: submissions/kuairand-pure-orpo-p5plus80-s1/. Reproduce: docs/reproduce_freeze.md.


Challenges we ran into

Date splits are easy to mis-describe as user-disjoint; they are not. A 7k user subsample with early-stop looked like an LLM vs BO bake-off but SFT/ORPO never left official FM—we do not mix those numbers with the 80-trial freeze.

HPUCB front-loads the LLM then drops ω_llm to ~0.1; after trial 50 it pulled only BO, so A5/T50/X50 never fired live.

SFT/ORPO were trained on --fast BO cards (epochs=8). Live search uses 20/30/40. SFT then proposed ep=8 on 20/21 LLM trials; ORPO collapsed to 59% duplicate configs. The hybrid still landed in the same basin as 80-BO.

We did not match our ML-track training budget (S1B: 7.5 SFT epochs, ~1.6k ORPO pairs, HP sweep). KuaiRand was 3 SFT epochs and 45 ORPO steps.


Accomplishments that we're proud of

A closed Piece-1 loop with kit-valid CSV, test unused for selection, and a matched 80-trial GP-EI control so we do not over-claim the LLM.

Modest, reproducible lift: official test 0.5946 → submitted 0.5969. ORPO LoRA was a real HPO arm (beat official on valid) while GP-EI found the shipping config.

Protocol discipline: frozen official 5 fields, no test in LoRA, OPE not used for selection, 7k/1K kept off the leaderboard claim.


What we learned

On this FM space, GP-EI with a full 80-trial budget already finds the basin. A LoRA proposer helps most if it diversifies early and BO refines; HPUCB that abandons the LLM before trial 50 does not let the prompt’s exploit tools run.

Training data beats extra epochs on the wrong teacher. Imitating fast-BO (ep=8) shows up at inference. Preference tuning (ORPO) moved proposals closer to the full-epoch basin than SFT, but did not beat a dedicated 80-BO run by a meaningful margin at n=1.

Official FM is a high bar; k barely moves primary. Honest reporting of “hybrid ≈ BO, LLM did not own the incumbent” is more useful than a larger-looking table.


What's next for recagent

Training on different takss or different splits can make the model better, with more data there would be more ways whcih we could split the data and provide more training data for the model.

Due to limited time we used GP-EI as the default BO algorithm what we can do is to try other BO algorithms and find the best one

In this way we could also build a similar task bank and have a GP trained on the different tasks top find similar tasks and use some methods like fanova to see if the run can be used, for those similar tasks we can add it into the prompt and provide a warm start for the LLM.

with more training data we can slowly update and train the model to evolve, this can create a self evolving framework where the model would get better and better, the model could be trained with updated trianing exmaple after k rounds of new tasks.

Retrain SFT→ORPO on full-epoch 80-trial traces (including 80-BO). Keep a LLM pull floor after trial 50 (or a 20-LLM → 60-BO phase) so A5 can fire; reject duplicate LLM configs.

Extra seeds. Flush token/GPU logs. Unfrozen FE/CASH or DeepFM only as a separate claim. 1K/27K only after a proposer trained for that scale—not as this freeze.

Built With

Share this project:

Updates

Submission history