Why Embedflix Wins

The 5 points that matter

# What we did Why it matters
1. Found the hidden metric constraint GAUC and nDCG@5 rank each user's own ~5 impressions. Any score term that is constant across those impressions is mathematically invisible to the metric. This tells the agent to prioritize item-side signals and user×item interactions, instead of wasting iterations on user-only features. It also explains the organizers' own ablation, where adding user/categorical features hurt performance.
2. Made the measurement trustworthy We discovered that the repository's claimed baseline was not actually the shipped FM. An unvalidated loss edit had replaced it and produced primary: None. We restored and pinned the official FM, reproducing 0.6016 ± 0.0003 across 5 seeds. Every subsequent improvement is measured against a verified baseline. We also replaced a noise-sensitive acceptance threshold with a two-stage seed-confirmed gate + no-op detector.
3. Built a metric-aligned model We use the FM logit as a feature in a LightGBM LambdaRank re-ranker, which directly optimizes ranking quality at @5. Features include leakage-safe train-only target encodings, engagement ratios, and user×item crosses. Results are rank-ensembled across 5 seeds. The model is optimized for the actual competition metric, rather than treating the FM score as the final prediction.
4. Made the agent genuinely autonomous Deterministic edits require 0 LLM calls. Harder code changes go through a code_writer with syntax validation, salvage, and corrective retries. Failed experiments trigger checkpoint/restore through error_recovery. The agent can run the full propose → edit → train → evaluate → reflect → recover loop without manual intervention.
5. Kept it practical NumPy + LightGBM, 0 GPU-hours, single CPU core, roughly 1–3 minutes per iteration. Haiku 4.5 handles inexpensive specialist reasoning; Opus 5 is reserved for code-writing. The system is cheap enough to actually run repeatedly, rather than relying on expensive GPU-scale experimentation.

Inspiration & Problem

Machine-learning engineers repeatedly run the same loop:

change code → train → evaluate → understand the result → try again.

That loop is structured enough to automate.

Track 2 asks us to build an agent that autonomously improves a ranking system on the KuaiRand-Pure short-video recommendation benchmark, beating the official Factorization Machine (FM) baseline.

The difficult part is not executing training. It is deciding what is worth changing, determining whether the change actually helped, and recovering when an experiment fails.

That is what Embedflix is designed to do.


What It Does

Embedflix is a LangGraph-based autonomous ML experimentation agent composed of a supervisor, six specialist proposers, a reasoning judge, and seven execution nodes.

Its most important contribution is not a model architecture. It is a metric-level insight.

GAUC and nDCG@5 rank impressions within each individual user's ~5-item list. Therefore, any score component that shifts every item for that user by the same amount cannot change the ranking and is invisible to the metric.

This has a direct consequence:

User-only features are often unable to improve the ranking. Item-side information and user×item interactions are much more valuable.

This explains the organizers' own ablation: adding user/categorical features reduced the score from 0.5953 → 0.5936.

We encode this insight directly into the supervisor and specialist prompts, preventing the agent from repeatedly proposing changes that the metric itself cannot reward.

Acting on the insight

Embedflix then:

  1. Restored the official baseline
    The repository's claimed baseline had been replaced by an unvalidated loss edit that never produced a valid score. We restored the shipped FM and pinned it as baseline_official.py.

  2. Made evaluation reliable
    tests/test_baseline_reproduces.py verifies byte identity and reproduces 0.6016 ± 0.0003 across 5 seeds. The agent now uses a two-stage seed-confirmed acceptance gate and explicitly detects no-op experiments.

  3. Changed the ranking architecture
    We use the FM logit as a feature inside a LightGBM LambdaRank re-ranker, with eval_at=[5], leakage-safe train-only target encodings, engagement ratios, and user×item crosses.

  4. Validated rather than cherry-picking
    The final result is rank-ensembled across 5 seeds to reduce variance. We also document rejected experiments, including 4-fold OOF stacking of fm_score, which reached only 0.5905 and overfit.

Result

Validation: 0.6045

Official baseline: 0.6016

Improvement: +0.0029

Test: 0.5975 (+0.0029)

And the entire pipeline runs on one CPU core with 0 GPU-hours.


How We Built It

1. Autonomous orchestration

Embedflix uses a LangGraph state graph:

Supervisor → Specialist → Judge → Code Writer → Train → Evaluate → Reflect → Repeat

When the required edit is deterministic, the agent takes a 0-LLM fast path.

For more complex changes, code_writer generates an exact OLD_CODE / NEW_CODE replacement. The pipeline then applies:

  • syntax validation
  • trailing-junk salvage
  • corrective retries
  • checkpoint/restore on failure

This means an experiment can fail without taking down the entire search.

2. Measurement integrity

Before searching for improvements, Embedflix verifies that the baseline itself is trustworthy.

The pipeline:

  • restores the shipped FM from git;
  • freezes it as baseline_official.py;
  • reproduces 0.6016 ± 0.0003 over 5 seeds;
  • screens candidate changes on seed 0;
  • confirms promising candidates across 3 seeds;
  • rejects bit-identical outputs as no-ops, rather than incorrectly labeling them failed techniques.

3. Metric-aligned re-ranking

Our final model uses:

FM logit → LightGBM LambdaRank → final ranking

Features include:

  • fm_score
  • train-only smoothed video target encoding
  • train-only author target encoding
  • five engagement ratios
  • user×item interaction features

Predictions are generated in data.load() row order to preserve row_id alignment, with dedicated alignment tests covering 19 checks.

We also tried 4-fold OOF stacking of fm_score. It scored 0.5905, showed clear overfitting, and was reverted. The failed approach and reason are preserved in the source rather than hidden.


Challenges We Ran Into

A contaminated baseline

The repository claimed a 0.6039 baseline, but we could not reproduce it from the shipped code. The associated run log contained primary: None.

Instead of optimizing against an unreliable number, we stopped and reconstructed the official FM first.

That changed the entire direction of the project.

Edits that silently did nothing

Our code_writer initially applied single-hunk find-and-replace edits. Some proposed changes technically succeeded but introduced code that was never actually used.

The result was an exactly identical score.

We turned this failure mode into a feature: Embedflix now detects no-op experiments and records them explicitly.

Seed noise

An initial 1e-4 improvement threshold accepted changes that were simply random seed variation.

We replaced it with a two-stage seed-confirmed acceptance gate, forcing promising changes to survive additional seeds before being accepted.


Accomplishments We're Proud Of

  • A metric-level insight that predicts why the organizers' own user-feature ablation failed.
  • A verified +0.0029 improvement over the official FM baseline.
  • A seed-confirmed result, rather than a cherry-picked best run.
  • 0 GPU-hours and single-CPU execution.
  • An autonomous loop with error recovery, checkpoint/restore, no-op detection, and deterministic fast paths.
  • Explicit documentation of rejected experiments, not just successful ones.

What We Learned

The most important lesson was that model performance starts before the model.

Understanding exactly what the evaluation metric rewards is often more valuable than adding another sophisticated architecture.

And before an agent can optimize anything, it needs to know that its measurements are trustworthy.

A weaker model with reliable evaluation beats a stronger model whose improvements you cannot trust.


What's Next

Three directions are already specified:

  • Causal user-session / sequence features — capturing what the user just watched, which the current FM + ID features cannot represent.
  • Matched-strength OOF stacking for fm_score.
  • Evaluation on the KuaiRand-1k and KuaiRand-27k bonus benchmarks.

The goal is not just to build an agent that can change ML code.

It is to build one that can understand what the metric rewards, form a hypothesis, run the experiment, reject noise, recover from failure, and learn from the result.

Built With

Share this project:

Updates

Submission history