Inspiration

Improving a recommender system usually involves the same slow loop: form an idea, change the code, run an experiment, study
the results, and decide what to try next. We built **EBK** to see how much of that process an AI agent could manage
independently while still making careful, transparent decisions.

## What it does

**EBK** is an autonomous research agent that improves recommender systems. It reviews earlier experiments, proposes a
hypothesis, modifies the pipeline, evaluates the result using **GAUC** and nDCG@5, and decides whether the change should be
accepted, rejected, or tested further.

Every experiment records what changed, why it was attempted, how the metrics moved, and how any failures were handled.

## How we built it

We built **EBK** in Python using pandas, NumPy, scikit-learn, LightGBM, PyTorch, Optuna, and Pydantic. Claude models support
experiment planning, reflection, and targeted code repairs, while Codex, VS Code, GitHub Actions, pytest, Ruff, and pre-
commit support development and validation.

We also added evaluation manifests, leak checks, preflight validation, ablations, promotion gates, and automated reports to
keep every result reproducible and trustworthy.

## Challenges we ran into

The hardest part was not getting the agent to suggest ideas. It was teaching it to judge those ideas carefully.

Small metric improvements can come from noise, inconsistent candidate sets, or data leakage. Experiments can also fail
halfway through a long run. We addressed these issues with confirmation tests, candidate-count diagnostics, recovery logic,
leak canaries, and strike counters.

## Accomplishments that we're proud of

**EBK** completed a converged run in 12 iterations, with no manual interventions and no **GPU** usage. The accepted checkpoint
achieved a KuaiRand-Pure validation score of 0.**601684**.

We are especially proud that **EBK** rejected a promising but insufficiently confirmed improvement. It showed that the agent
could prioritize trustworthy evidence over simply selecting the highest observed score.

## What we learned

We learned that autonomous research involves much more than tuning a model. A useful research agent must design meaningful
experiments, evaluate them consistently, recover from failures, and know when the evidence is strong enough to trust.

We also learned that failed and rejected experiments remain valuable when their reasoning and results are properly
recorded.

## What's next for EBK

Next, we want to improve automated feature discovery, strengthen uncertainty-aware promotion decisions, run experiments in
parallel, and evaluate **EBK** on the KuaiRand-1k and KuaiRand-27k bonus benchmarks.

Our longer-term goal is to help researchers spend less time managing experiments and more time asking better questions.

Built With

+ 3 more
Share this project:

Updates

Submission history