Inspiration
We wanted to explore whether an AI agent could do more than just generate code. A lot of ML work is actually very repetitive: try an idea, change the model, run it, check the metrics, figure out what went wrong, and try again.
So we thought it would be interesting to build an agent that could go through that process by itself and actually learn from what happened in previous iterations.
What it does
Salty Crocodile is basically an autonomous research agent for recommender systems.
It looks at the current best model and past experiment results, decides what to try next, gets an LLM to write the code for that experiment, runs it, evaluates it, and then decides whether the new version is actually better.
If something breaks, the agent can also retry, debug the same idea, or fall back to the last working model instead of just stopping.
The main idea is that the LLM is not making every decision. Our agent controls the overall research process and uses the LLM mainly to implement the experiments.
How we built it
We built the project in Python using the KuaiRand-Pure starter kit.
We started with the provided Factorization Machine baseline and built an agent loop around it. The agent can explore different directions like changing the loss function, using user history, trying multi-task signals, modelling watch time, or looking at time-related effects.
For code generation, we used NVIDIA Nemotron-3 Super 120B through the Swiss AI API.
The rest of the system handles things like choosing what experiment to run, keeping track of past results, checking whether an improvement is actually meaningful, handling failures, and deciding which model to keep.
We also made sure the agent only sees the training and validation data, so it cannot accidentally use the hidden test set while experimenting.
Challenges we ran into
The biggest challenge was getting reliable code from the LLM.
Sometimes it would generate incomplete code, hit the output limit, or make a small mistake that caused the whole experiment to crash. Pairwise ranking was especially annoying because indexing and gradient-shape bugs were very easy to introduce.
We also realised that a slightly higher score does not always mean the model is actually better. The baseline itself has some variation across runs, so we had to avoid accepting tiny improvements that could just be noise.
Another challenge was deciding what the agent should try next. We only have a limited number of iterations, so it cannot just test everything.
Accomplishments that we're proud of
We are probably most proud that the project actually behaves like an agent rather than just a wrapper around an LLM.
It can choose experiments, react when something fails, keep the best working model, and use previous results when deciding what to do next.
One of our runs found a pairwise-ranking approach that reached:
- GAUC: 0.6702
- nDCG@5: 0.5372
- Primary: 0.6037
The official validation baseline was 0.6016.
It was also nice to see the agent recover from failed attempts and eventually produce a working improvement instead of needing us to manually fix the generated model during the run.
What we learned
One thing we learned very quickly is that building an agent is not just about having a good prompt.
A lot of the important parts are outside the LLM: what information you give it, how you choose the next experiment, how you handle failures, and how you decide whether a result is actually worth keeping.
We also found that making the model bigger was not necessarily the best direction. Changing the objective to better match the ranking metrics ended up being more useful in our experiments.
And weirdly, failed runs were still useful because they gave the agent information about what not to repeat or what needed to be fixed.
What's next for Salty Crocodile
If we had more time, we would mainly make the agent more reliable and give it better memory of what it has already tried.
We would also explore better negative sampling for pairwise ranking, combining ranking loss with the original pointwise objective, user-history models, multi-task learning, and better watch-time modelling.
We would also add stronger duplicate detection, because sometimes the agent can spend another iteration producing something that is basically the same as the current best model.
Longer term, we think the same setup could be useful outside recommender systems too: give the agent a model, a metric, and an experiment budget, and let it work through possible improvements on its own.
Log in or sign up for Devpost to join the conversation.