-
Agent AI-yoh is an auditable autonomous recommender research agent, abbreviated as “Agent A” in the repository.
-
Agent AI-yoh turns user, video, and logged-feedback data into one score per exposure, ranking each user’s short-video feed.
-
Agent AI-yoh turns manual ML iteration into a bounded loop that proposes, trains, validates, learns, and locks the best model.
-
Safe data, unified candidates, validation-only selection, persistent ledgers, recovery, and immutable locks form the architecture.
-
Six controlled anchors led to three blend trials, with trial-07 locking a 65% FM and 35% LambdaRank ensemble.
-
Validation reached 0.66910648 GAUC, 0.53661937 nDCG@5, and 0.60286292 primary using 9/50 trials and zero test metrics.
-
The agent moves from hypothesis to controlled experiment, validation, diagnosis, and the next evidence-backed experiment.
-
Next: independent benchmarks, cross-seed checks, richer causal models, faster training, and a more capable auditable agent.
Inspiration
Machine-learning teams spend enormous effort repeating the same cycle: forming hypotheses, building models, running experiments, interpreting results, and deciding what to try next. In recommender systems, meaningful progress often depends on many small, carefully controlled improvements.
AI-yoh! was created around a simple idea: research agents should help teams explore this search space more efficiently while remaining transparent and accountable. Agent AI-yoh is our first step toward that vision—a system that can run structured experiments, learn from evidence, work within clear constraints, and leave behind a record that people can review and reproduce.
What it does
Agent AI-yoh manages the recommender-system research loop within a fixed experimental budget. It evaluates baseline and follow-up ideas, uses validation evidence to guide the next step, recovers from failed trials, and stops when additional exploration is unlikely to provide meaningful gains.
The platform supports several modeling approaches, including Factorization Machines, ListNet, causal history features, BPR, multi-task objectives, LambdaRank, and model ensembles. Every experiment is tracked, including failed and interrupted runs. At the end of the process, the system locks the best validation checkpoint and generates one valid ranking score for each test exposure without accessing test labels.
How we built it
We built Agent AI-yoh as a modular Python research platform using NumPy, LightGBM, Optuna, SciPy, SQLite, and OpenAI Codex. A dedicated safety layer protects the temporal data boundary and ensures that model selection is based only on permitted validation evidence.
Each dataset receives its own research ledger, experiment budget, storage, and model lock. Experiments are represented by canonical configurations, allowing identical work to reuse existing evidence. The system also maintains an append-only SHA-256 decision chain containing hypotheses, evidence, alternatives, selected actions, metrics, code provenance, and recovery events.
Automated tests cover leakage prevention, budget enforcement, checkpoint recovery, gradient behavior, score alignment, and validation-only selection.
Challenges we ran into
Our main challenge was improving a strong Factorization Machine baseline without violating the benchmark’s temporal structure. Historical and behavioral features had to be generated causally, using only information available before each event. Validation and test features were frozen from training data, with explicit safeguards against current-row, future, and label leakage.
We also learned that complexity does not guarantee progress. Several sophisticated approaches did not outperform the baseline. Rather than treating those results as failures to hide, Agent AI-yoh retained them as evidence and used them to focus the remaining search.
Accomplishments
Agent AI-yoh completed the formal research run in 9 of the permitted 50 iterations and automatically selected the best validation checkpoint. The final system combines the official Factorization Machine with a causal-behavioral LambdaRank model.
It achieved a validation primary score of 0.60286292, compared with the organizer’s rounded baseline of 0.6016. This is a modest positive improvement, below the declared significance threshold, so we do not present it as a statistically significant result.
Beyond the metric, we are proud of the research infrastructure: the final prediction file was generated without test-label access, and the model, predictions, experiment history, and decision records can be verified through reproducible hashes and audit artifacts.
What we learned
The strongest lesson was that better research decisions often matter more than a more complicated model. LambdaRank was weaker on its own, but its errors were sufficiently different from the Factorization Machine to provide a small ensemble benefit.
We also found that reliable recommender research depends on disciplined data construction, correct grouping, score normalization, checkpoint selection, and strict evaluation boundaries.
Most importantly, autonomous ML systems need governance as well as intelligence. A useful research agent must know how to explore, when to stop, how to handle failure, and how to explain the path it took.
What’s next for AI-yoh!
Our next step is to evaluate Agent AI-yoh across independent KuaiRand benchmark fingerprints, transferring research methods rather than datasets, checkpoints, or conclusions. We also plan to improve feature efficiency, add cross-seed finalist verification, and expand the agent’s ability to design new candidates within a validated safety contract.
Our broader vision is practical: help ML teams spend less time managing repetitive experimentation and more time making high-quality research decisions. We believe progress in this direction will come from combining autonomous exploration with strong evaluation boundaries, reproducibility, and human oversight.
AI-yoh! is not a claim that research can be fully automated today. It is an experiment in building research systems that are more systematic, efficient, and trustworthy.
Built With
- aigagent
- artificialintelligence
- automl
- autonomousagent
- dataleakageprevention
- experimenttracking
- explainableai
- factorizationmachines
- kuairand
- lambdarank
- learningrank
- lightgbm
- listnet
- machine-learning
- mlops
- numpy
- openaicodex
- optuna
- python
- rankingsystems
- recommendersystems
- reproducibleresearch
- responsibleai
Log in or sign up for Devpost to join the conversation.