Inspiration
Robot-learning researchers build AI models that control robot arms, and they test them in a shared simulator. On one task, the robot simulator these researchers treat as standard (SIMPLER) said a model fails 96% of the time. On a real robot arm, humans watched that same model succeed 92% of the time.
Both numbers are from AutoEval, and if you believed the simulator, you would have thrown away a model that actually succeeds 92% of the time.
Testing on the real arm instead is slow and expensive. The best public comparison of these models used 1,500 real robot attempts to produce 30 success rates, six models on five tasks, with one dedicated arm, a person watching and grading every attempt, and twenty-minute pauses so the motors could cool. That is why these researchers compare six models and not six hundred.
Three research groups have already replaced the robot with a world model: WorldGym, WorldEval, and dWorldEval. Their main question is whether evaluation in the world model agrees with evaluation on the real robot.
We asked a different question: if nothing changes, do you get the same ranking twice?
Imagine a bathroom scale that reads 70 kg, then 74 kg, then 66 kg for the same person. Its average might be correct, but it is useless for measuring a 1 kg change because its own variation is larger than the effect you care about.
A policy evaluator has the same problem. If its run-to-run variation is larger than the improvement between two policies, you cannot reliably tell whether a change helped or hurt.
So before using Kosmos to rank policies, we measured the evaluator itself.
What it does
Kosmos ranks the VLA models that drive a robot arm, without using a robot.
Those models are called policies. A policy takes a camera image plus a task written in English, and outputs how the arm should move next. We test three open ones: OpenVLA, Octo, and MiniVLA.
The loop:
- You type a task, for example "pick up the bowl".
- The policy looks at the current scene and outputs a movement.
- Instead of sending that movement to hardware, we send it to a world model: an AI trained on real robot video that draws the next frames given the current frame and the commanded movement. We use NVIDIA Cosmos.
- We feed the last generated frame back to the policy and repeat. After a few rounds we have a video of an attempt that never physically happened.
- A scorer decides whether the task was completed.
That produces a ranking. A ranking by itself proves nothing, because three groups have already published one. So we tested our own system the way you would test that bathroom scale, asking not "is the number right?" but "how much does the number move when nothing about the models has changed?"
Four questions about the measuring tool itself. We could not find any of them answered in published work:
- Does it repeat? Split the attempts into two halves and check whether both halves rank the models in the same order.
- How small a difference can it detect? Below some gap, two models cannot be told apart inside our own noise. We report where that gap is.
- How long before it stops being accurate? Each generated frame is built on the previous generated frame, so errors accumulate. We measured how many steps we get before the ranking breaks.
- How cheap can it get? Changing only the request settings moves the cost of a single attempt by more than ten times. We found the cheapest setting that still ranks correctly.
We also report the conventional check, so we can be compared directly against published work: how closely our per-task success rates match AutoEval's 1,500 human-graded real attempts.
How we built it
Kosmos is a heterogeneous inference pipeline running on Baseten.
Chains: Policy inference, world-model generation, deterministic validation, and VLM judging are separate stages with very different compute profiles. The world model is the dominant GPU workload, the policy has its own model-residency requirements, the validity gate is CPU-side Python, and the judge is a separate multimodal model. Keeping them separate lets each stage use its own hardware profile and scale independently instead of provisioning the entire pipeline for its most expensive component.
Diffusion-aware serving: The world model runs as a dedicated warm deployment with
concurrency_target = 1. A single video-generation request can already occupy most of a GPU, so stacking independent generations on one replica mostly creates compute and memory contention. Throughput instead comes from horizontal replica scaling and batching where appropriate.Asynchronous rollout fan-out: Different robot rollouts are independent, so evaluations can be parallelized across the serving fleet. We submit rollout work asynchronously, allowing many requests to remain in flight while Baseten queues them according to available capacity. Submission and result retrieval are decoupled, and the frontend exposes queue depth and completion state while an evaluation is running.
Large-scale evaluation serving: Kosmos queues and runs rollout batches across concurrent world-model workers, then sends completed trajectories through a separate VLM judging stage. We support configurable worker concurrency, judge batching, retries, and campaign-level aggregation, so experiments like 4 policies × 25 tasks × 4 seeds = 400 rollouts can run as one evaluation instead of trajectory by trajectory.
Warm model residency: Policy and world-model workers load their weights once and reuse them across rollout steps. We therefore distinguish cold-start/model-load time from warm inference latency instead of treating them as one latency number.
Self-distilled VLM judge: Scoring thousands of generated rollouts with a large multimodal model quickly becomes the dominant evaluation cost. We use the judge itself to generate supervision over rollout videos, then post-train a smaller model from the same model family on those judgments using Baseten Training Jobs. This gives us a cheaper specialized evaluator for high-volume scoring while keeping the grading objective fixed. Before using the distilled judge for policy evaluation, we measure how closely its judgments match the original teacher.
Training Jobs: We use Baseten Training Jobs to post-train a smaller multimodal judge from our larger evaluator. This moves the high-volume scoring path onto a cheaper model while attempting to preserve agreement with the original judging protocol.
Before paying for semantic VLM evaluation, we also run a deterministic action-consistency gate.
The policy outputs a commanded robot action, while the world model generates what should happen visually after that action. We compare measurable motion in the generated trajectory against the commanded action and reject obvious conditioning failures before they reach the judge. This prevents visually plausible but action-inconsistent world-model generations from contaminating the evaluation.
We optimize for total rollout throughput rather than the latency of any single request. We track:
- queue depth
- active and desired replicas
- warm inference latency
- GPU-seconds per rollout
- rollouts per GPU-hour
- cost per evaluation batch
Because independent rollouts can be distributed across independent workers, Kosmos scales out across replicas rather than distributing a single rollout across multiple GPUs. This lets large evaluation runs burst across the serving fleet and scale back down once the queue clears.
We measure how far the gripper shifts between the generated frames and compare it to the commanded movement. If the two numbers are a mismatch, then we discard that video before scoring. Baseten's engineers made the same call: they replaced an LLM that was "essentially pattern-matching against rules" with Python, and reported "improving accuracy to 100%" (emergency medicine case study).
Challenges we ran into
- Our first design could not support its own claim: a single correlation across five models had a confidence range including zero, so we now compare each policy on each task, which gives five times the sample.
- A nearly frozen scene beats every automatic video-quality score: that is why the arithmetic check runs before the scorer, and we confirmed it works by feeding in shuffled commands and watching the success rates collapse.
- Discarding unusable video biases the ranking: we report both the raw rate and a worst-case range whose width is exactly equal to the discard rate.
- Our own cost estimate was wrong: we priced the run as GPU time and forgot the scorer, which turns out to dominate the bill.
Accomplishments that we're proud of
- Audited our ground truth: AutoEval and SIMPLER disagree sharply on Octo, so we report results both with and without it.
- Tested whether training loss predicts scorer quality and found that it does not reliably capture evaluator behavior.
- Built a more explicit grading protocol than the source data provides, including multiple graders, blinding, and inter-grader agreement.
What we learned
- Building the evaluator is the easy part, and proving that it measures anything is not: Baseten's own eval paper makes the same argument, and we spent more of the weekend on that proof than on the evaluator itself.
- The standard ranking-error measure has two published definitions that disagree on identical data: we name the one we use and report what a random guess would score, so the number has a floor.
What's next for Kosmos
- Publish it as a versioned environment: OpenEnv already hosts environments, but nobody versions benchmark results, so today a stranger cannot rerun our evaluation and arrive at our number.
- Write down the contract for how an evaluation harness talks to a robot policy: every standard today assumes it is calling a chat endpoint, whereas a policy takes images and returns motion.
- Test a policy from an unrelated research lineage: ours and our world model come from overlapping research groups, which plausibly makes our numbers look better than they should, and we have not measured how much.
Log in or sign up for Devpost to join the conversation.