Team Toasted Presents ToastBot

Inspiration

We believe the best innovations come from a human-centered design process. So, when looking for project ideas, we were inspired by Carolyn, a chef who shares disability-friendly recipes on TikTok as @epicuriousexpeditions.

Her journey started when she injured her knee and had surgery. Her recovery impacted her ability to cook, which led her to creating and sharing her disability-friendly recipes on the internet. Learning about her situation, and the situations of people like her, made it clear that this was a real opportunity to build something useful. Cooking is something most people take for granted, but for someone who can't safely use a knife, a stovetop, or a hot oven, it can become an obstacle.

Our Design Process

We followed a standard design process, and we also tried to hold ourselves to a design ethos throughout. This ensured that our project was ethically designed and effective for user needs.

Research

For user research, we gathered real accounts from people cooking with disabilities, including multiple YouTube videos. This showed us firsthand how much extra planning and risk goes into something as simple as making a meal.

For contextual research, we also looked at what's already on the market for people who need safer cooking options, like Instant Pots and slow cookers.

What we noticed is that these existing appliances solve part of the problem, but not all of it. They're safer than a stovetop, but they are limited to cooking in one pot with only a few settings available.

Define

Bringing our research together, we mapped out the causes and effects of the core problem we kept running into: disability-friendly meals can taste boring and repetitive.

We realized that the core issue on hand is that safety issues prevent users from using high-heat appliances to make crispy foods. For example, users with tremors, limited range of motion, and endurance issues all pointed out how it is simply too unsafe to interact with high heat, forcing them to get that texture from prepared foods.

Ideate

From there, we landed on our guiding question: how might we use a robot arm to make crispy food for individuals who cannot safely use high-heat appliances?

We chose to use the robot arm to make a toasted peanut butter and jelly sandwich. PB&J's have both crispy texture from the toast and smooth texture from the filling, allowing us to target the core problem of no variety. It also acts well as a baseline to evaluate this projects viability, and as a childhood favorite, it allows us to empathize with our users on a greater level.

Build & Test

Hardware & Materials

We built ToastBot around a Seeed reBot B601-RS, a 7-DoF arm, paired with a matching leader arm for teleoperation. There are two RGB cameras: one mounted overhead for the whole scene, and one on the wrist so it can see what it's actually grabbing.

Hardware and materials

The rest of the setup resembles a normal kitchen, including bread, peanut butter, jelly, a toaster, a pan, and a spatula. We are also proud to say we made a few 3D-printed tools, including a bread holder, a gripper-friendly spatula, and a toast rack. These were designed so the arm's gripper could handle them reliably.

Our pipeline

End-to-end pipeline

  1. Teleoperate the robot. We drive the leader arm by hand, and the follower arm mirrors us, recording two camera views plus every joint position throughout the episode
  2. Build the dataset. Those recordings become a LeRobot Dataset made by RGB images paired with the exact robot actions we took.
  3. Train the policy. A frozen DINOv2 visual encoder turns the camera images into features, and a Flow Matching policy learns to turn those features into motion.
  4. Deploy. Load the trained checkpoint onto the robot and let it run: observe, predict actions, execute, repeat at 30 Hz.
  5. Execute the task. Move the bread to the toaster, pull the lever, and move the toast to the plate.

The key thing here is that we never wrote a single trajectory by hand. The robot learned the motions over about 40 trials per task, giving it the ability to adapt based on potential discrepancies in the setting.

Teleoperating the arm

How we approached it

Five decisions shaped the whole build, and most of them were about doing less, not more.

1. Small task-specific policies. With only 40–50 demonstrations per task, our bottleneck was data, not compute. So we froze the DINOv2 encoder rather than training it. That preserves the visual features it already learned from millions of images, cuts our training cost substantially, and keeps a model that big from simply memorizing our forty examples.

2. A state machine for the long-horizon task. The full sequence is place → pull lever → collect, with a long toaster wait in the middle. We could have tried to train one policy for all of it. Splitting it into three made failures easy to isolate, kept each episode short, and means we can add new stages later without retraining anything.

3. No VLA. Vision-language-action models are the exciting thing in robotics right now, and we deliberately didn't use one. For three fixed tasks, a large VLA would mostly be acting as an expensive instruction-to-policy lookup which wouldn't fit on an 8 GB Orin Nano.

4. Make the scene trainable. Because our cameras are rigidly mounted and our actions are absolute joint positions, where something appears in the image genuinely means something. So instead of aggressively randomizing the visuals, we went the other way with fixed camera placement, clean background, controlled lighting, and only mild augmentation.

5. Recording quality mattered Our first recordings were fixed 20-second windows, which meant every episode ended with several seconds of the arm sitting still. We fixed it by ending episodes the moment the task succeeded, trimming the idle frames out of what we'd already recorded, and collecting more data.

Testing

We evaluated each policy offline against replayed episodes before ever letting it drive the arm, then ran it live. Our strongest policy is the lever pull:

Metric Result
Step error 0.173 joint units — about 1% of the range of motion
Re-plan latency 16.5 ms
Observation conditioning 64.9x

That last number is the one we're happiest about. It measures how much the robot's output actually changes when you change what it sees. A high number means it's genuinely looking at the scene and reacting — not replaying a memorized motion with its eyes closed. That's the whole point of the project in a single statistic.

We also chose Flow Matching specifically to make this deployable. It's closely related to the diffusion models behind AI image generators, but it learns straighter paths, which means it can produce a motion in 5 steps instead of roughly 100.

Testing the arm

Challenges we ran into

A bug that made our images silently wrong. The image statistics our pipeline computed came out with a standard deviation of exactly zero — an integer overflow deep in the library was wrapping values around and collapsing the math. Nothing crashed. Nothing warned us. It would have fed the model wildly mis-scaled images forever. We found it because a value of precisely 0.0 looked too clean to be real, and patched it upstream.

A dataset labeled with the exact opposite of what it showed. Our recording script had a default task description, and one session silently inherited the previous task's label — so 40 demonstrations of putting bread in shipped labeled as taking toast out. The script now refuses to start without you saying what you're recording.

Bread is a genuinely hard object. It's deformable, it varies slice to slice, and it doesn't sit where you put it. A lot of our 3D printing was in service of giving the gripper something rigid and predictable to work with.

Hardware that fights back. We hit CAN bus buffer overflows that would stall recordings partway through, and had to harden the whole recording loop before we could reliably capture 40 episodes in a row.

Accomplishments that we're proud of

  • ~130 demonstrations we teleoperated by hand, across three tasks, cleaned and documented.
  • A robot that does a complete multi-stage kitchen task end to end — load, actuate, retrieve — entirely from learned behavior.
  • Sizing every decision for hardware a real person could afford, rather than a datacenter.
  • Finding and fixing a real bug in the underlying library that was producing silently wrong numbers.
  • Being honest about our own results. Every number above is measured, and we say clearly which parts we've tested on hardware and which we haven't yet.

What we learned

The model was never the hard part. The hard part was the data — and specifically the failures that don't announce themselves. A statistic of zero, a reversed label, a few seconds of idle time at the end of every recording. Each one produced a policy that trained fine, reported a perfectly reasonable loss, and then did the wrong thing on the real robot. We spent more time building tools to catch those than we did on architecture.

We also learned that constraints make better projects. Committing to a small, cheap device forced us toward Flow Matching, a frozen encoder, and three small policies instead of one big one — and that turned out to be a better system than the unconstrained version would have been.

What's next for ToastBot

  • The full sandwich. Right now we handle the toast. Spreading peanut butter and jelly is a genuinely harder manipulation problem, and it's the obvious next step.
  • More kitchen tasks. The pipeline is task-agnostic — teleoperate, record, train, deploy. Adding a new skill means collecting an afternoon of demonstrations, not writing new code.
  • Voice control. We built a language-conditioned policy but haven't deployed it. That gets us to "put the bread in" as something you say out loud.
  • Testing with the people we built this for. Everything so far has been us, in a lab. The next real milestone is putting it in front of someone who cooks the way Carolyn does and finding out what we got wrong.

Safety

The robot arm moves on its own during evaluation, including between tasks as it returns to its home position. We keep an emergency stop within reach for every run, and the software runs a full preflight check with the motors powered down first — because a problem you discover after the arm is live is a problem with a moving arm.


Built on LeRobot by Hugging Face, with a frozen DINOv2 visual encoder from Meta AI. Flow matching per Lipman et al., ICLR 2023.

Built With

Share this project:

Updates

Submission history