Inspiration

ML engineers spend most of their time doing the same loop over and over by looking at the data, try something, check the score, figure out why it didn't work, then try again. The Track 2 asked if an AI agent could do that loop by itself on a real dataset, without anyone telling it what to try. I thought it was a really cool challenge and wanted to build something where I could actually prove the agent was not cheating or getting hints, not just say it wasn't.

What it does

LongView is an agent that takes the KuaiRand-Pure video recommendation dataset and a scoring rule, then runs the whole ML research loop on its own. It comes up with an idea, writes the code, trains it, looks at its own results, and decides what to try next, over and over until it stops improving. It's made of 3 parts, a scoring system the agent can never touch or cheat on, a 2 model brain, one that is a strong model for ideas, one is a cheaper faster model for writing and fixing code, as well as a search process that double checks every result with multiple random seeds before trusting it. It ended up managing to beat the baseline on the hidden test set, using 24 of the allowed 50 tries, which costs $6 in API costs, under 3 hours, and zero times where I had to step in and help it.

How we built it

The main idea was to make the no cheating rules actually impossible to break instead of just asking the agent "nicely" not to break them. Every model the agent writes runs in its own separate process, and that process is given data where the answers (test labels) are already deleted before it even starts, so there's nothing to leak. When it wants to check how well it's doing, it calls a function that gives it back just a score, not the actual answers. Everything is written in plain numpy since that's what the competition required. I used one AI model to plan and write hypotheses and a cheaper, faster one to actually write and fix the code. I also copied the organiser's scoring code exactly and hashed it so I would know immediately if anything got changed on accident.

Challenges we ran into

My first full run finished too fast, right at the earliest checkpoint I allowed it to stop, and the result looked good at first. But when I checked the math properly, the "improvement" was basically just random noise (approximately like 70% chance of being nothing). So instead of just accepting that, I went ahead and read the actual code the agent had written for its two most ambitious ideas, then I found that it had bugs in it. One used a single previous video ID as a feature, which is extremely specific to ever repeat for the same user, so the model could not learn anything from it. Another bug was that it had swapped out a normal loss function for something that did not actually match how it was being graded. So, I rebuilt the search process to fix these kinds of problems in general, without literally telling the agent the fix which is akin to giving it the answer key directly. Now every result gets checked 3 times with different random seeds before it counts. Ideas that don't quite work still get saved so they can be tried again later with more info instead of just getting thrown away. I also ran into random numpy bugs such as where scores were silently getting corrupted from an integer overflow, some annoying Python import issues, and at one point I thought the whole run had frozen for over an hour when it was actually just a really slow API response, so I had to add live logging just to tell the difference.

Accomplishments that we're proud of

The biggest one is that no cheating system is not just a rule, but it is enforced into the code so it is impossible to break, and this can be proven from the logs. The real scored run had 0 intervention from me, and the agent did everything on its own. Instead of just keeping the fluke result from the first run because it looked fine, I managed to identify it, fix the actual problem, and re-ran it to get a smaller but real result. In addition, the agent figured out on its own two of the exact same research directions TikTok said were the most promising ones, without ever being told that.

What we learned

One result from 1 random seed does not mean much. It is basically a coin flip, and it can be easy to accidentally build a system that just gets lucky instead of actually improving. I also learnt that AI API response times can vary a lot more than you would expect, some takes 10 minutes, some take 2 minutes, hence proper logging is crucial if not you will have no idea if something broke or is just slow. Probably the biggest thing was realising how tempting it is to just directly tell the agent the fix once you figure it out. Not doing that was honestly harder than build the actual safety sandbox, but it's the only way the agent's results actually mean something.

What's next for LongView

The next step is giving the agent better ways to represent user history, like using rolling stats, (e.g. how long they watch videos, how recently active, etc.) instead of raw IDs, since that is why its sequence modeling attempt did not work. We would also want to replace our simplified way of picking which idea to improve next with something much smarter, always double checking results with multiple seeds even early on, or maybe letting the agent keep some memory of what it has learned across different runs instead of starting fresh.

Built With

Share this project:

Updates

Submission history