Inspiration

Machine-learning teams spend enormous effort repeating the same cycle: forming hypotheses, building models, running experiments, interpreting results, and deciding what to try next. In recommender systems, meaningful progress often depends on many small, carefully controlled improvements.

AI-yoh! was created around a simple idea: research agents should help teams explore this search space more efficiently while remaining transparent and accountable. Agent AI-yoh is our first step toward that vision—a system that can run structured experiments, learn from evidence, work within clear constraints, and leave behind a record that people can review and reproduce.

What it does

Agent AI-yoh manages the recommender-system research loop within a fixed experimental budget. It evaluates baseline and follow-up ideas, uses validation evidence to guide the next step, recovers from failed trials, and stops when additional exploration is unlikely to provide meaningful gains.

The platform supports several modeling approaches, including Factorization Machines, ListNet, causal history features, BPR, multi-task objectives, LambdaRank, and model ensembles. Every experiment is tracked, including failed and interrupted runs. At the end of the process, the system locks the best validation checkpoint and generates one valid ranking score for each test exposure without accessing test labels.

How we built it

We built Agent AI-yoh as a modular Python research platform using NumPy, LightGBM, Optuna, SciPy, SQLite, and OpenAI Codex. A dedicated safety layer protects the temporal data boundary and ensures that model selection is based only on permitted validation evidence.

Each dataset receives its own research ledger, experiment budget, storage, and model lock. Experiments are represented by canonical configurations, allowing identical work to reuse existing evidence. The system also maintains an append-only SHA-256 decision chain containing hypotheses, evidence, alternatives, selected actions, metrics, code provenance, and recovery events.

Automated tests cover leakage prevention, budget enforcement, checkpoint recovery, gradient behavior, score alignment, and validation-only selection.

Challenges we ran into

Our main challenge was improving a strong Factorization Machine baseline without violating the benchmark’s temporal structure. Historical and behavioral features had to be generated causally, using only information available before each event. Validation and test features were frozen from training data, with explicit safeguards against current-row, future, and label leakage.

We also learned that complexity does not guarantee progress. Several sophisticated approaches did not outperform the baseline. Rather than treating those results as failures to hide, Agent AI-yoh retained them as evidence and used them to focus the remaining search.

Accomplishments

Agent AI-yoh completed the formal research run in 9 of the permitted 50 iterations and automatically selected the best validation checkpoint. The final system combines the official Factorization Machine with a causal-behavioral LambdaRank model.

It achieved a validation primary score of 0.60286292, compared with the organizer’s rounded baseline of 0.6016. This is a modest positive improvement, below the declared significance threshold, so we do not present it as a statistically significant result.

Beyond the metric, we are proud of the research infrastructure: the final prediction file was generated without test-label access, and the model, predictions, experiment history, and decision records can be verified through reproducible hashes and audit artifacts.

What we learned

The strongest lesson was that better research decisions often matter more than a more complicated model. LambdaRank was weaker on its own, but its errors were sufficiently different from the Factorization Machine to provide a small ensemble benefit.

We also found that reliable recommender research depends on disciplined data construction, correct grouping, score normalization, checkpoint selection, and strict evaluation boundaries.

Most importantly, autonomous ML systems need governance as well as intelligence. A useful research agent must know how to explore, when to stop, how to handle failure, and how to explain the path it took.

What’s next for AI-yoh!

Our next step is to evaluate Agent AI-yoh across independent KuaiRand benchmark fingerprints, transferring research methods rather than datasets, checkpoints, or conclusions. We also plan to improve feature efficiency, add cross-seed finalist verification, and expand the agent’s ability to design new candidates within a validated safety contract.

Our broader vision is practical: help ML teams spend less time managing repetitive experimentation and more time making high-quality research decisions. We believe progress in this direction will come from combining autonomous exploration with strong evaluation boundaries, reproducibility, and human oversight.

AI-yoh! is not a claim that research can be fully automated today. It is an experiment in building research systems that are more systematic, efficient, and trustworthy.

Built With

  • aigagent
  • artificialintelligence
  • automl
  • autonomousagent
  • dataleakageprevention
  • experimenttracking
  • explainableai
  • factorizationmachines
  • kuairand
  • lambdarank
  • learningrank
  • lightgbm
  • listnet
  • machine-learning
  • mlops
  • numpy
  • openaicodex
  • optuna
  • python
  • rankingsystems
  • recommendersystems
  • reproducibleresearch
  • responsibleai
Share this project:

Updates

Submission history