Inspiration
In Interstellar, a giant wave hits the crew on Miller's planet and Brand ends up pinned under wreckage in the water. CASE folds its body into a wheel, rolls over to her, and carries her out.
We kept thinking about that scene because real rescuers face a smaller version of it. After the World Trade Center collapsed in 2001, teams sent small robots into gaps in the rubble that were too tight or too unstable for a person. Collapses keep creating the same problem: someone is trapped where rescuers can't safely go, and time is running out.
Here's the catch. A robot with a fixed body has to pick its shape on day one. Make it tall and it can step over a crack, but it won't fit under a fallen beam. Make it short and it slides under the beam, then gets stuck at the crack. A collapse site has both within a few meters of each other. CASE worked because it could change shape when the moment needed it. We wanted to see if a robot could learn to do that by itself.
The theme was Into the Skies, so we put the rescue on the Moon: a lunar base hit by a moonquake, with one astronaut stuck at the end of a broken corridor. Nobody else is coming.
What it does
Ariel is a simulated rescue robot that can change the length of its own limbs while it moves. Its thighs, shins, arms and torso all telescope, like a car antenna, so it can shrink to get low or stretch out to take a bigger step. It runs under both Moon gravity (1.62 m/s²) and Earth gravity (9.81 m/s²).
The course is a wrecked lunar base laid out like a collapse site:
- Crates block the path, so Ariel has to steer around them.
- A low deck hangs ahead, so it has to pull in its legs, torso and arms to get under.
- A gap in the floor takes longer shins to step across.
- A rocky slope leads up to the astronaut.
Nobody tells Ariel when to change shape, apart from a small hint early in training that fades out. It figured that out on its own through reinforcement learning, which is trial and error scored by a reward, over millions of attempts. We recorded every attempt, so you can scrub through training and watch early versions fall over and later ones duck under the deck.
How we built it
The body. Ariel starts from the CMU humanoid used in MoCapAct. We added nine telescoping segments, each one a tube that slides inside another. A segment can stretch between 80% and 140% of its normal length, and it can only change length at 0.25 m/s:
$$ 0.8 \ell_0 \le \ell \le 1.4 \ell_0, \qquad |\dot{\ell}| \le 0.25 \text{ m/s} $$
Here ℓ₀ is the normal length. The speed limit matters. Without it, Ariel snaps its legs out and launches itself like a catapult.
The walking. Teaching a humanoid to walk from scratch would have eaten our whole hackathon, and our first try proved it. So the walking comes from MoCapAct, a model Microsoft researchers trained on human motion-capture data. We keep it frozen and never change it.
The part we trained. On top of the walker sits our own policy, trained with PPO, a standard reinforcement learning algorithm. 33 times a second it decides how fast to go, how to steer, small corrections to the walking, and how long each group of limbs should be.
The reward. RL agents will exploit any loophole you leave. So Ariel only earns points for getting farther than it has ever been:
$$ r^{\text{progress}}_t = \max\Big(0, x_t - \max_{k \lt t} x_k\Big) $$
Lunging forward or rocking back and forth earns nothing. Changing shape has no reward of its own, but it costs energy, so Ariel only does it when it helps reach the astronaut.
The training wheels. Early on, two hand-written helpers give hints. One suggests a body shape for the deck and the gap, and the other suggests a path around the crates. Their influence shrinks to zero over training, and our final tests never use them.
The fair test. The layout is randomized every run so Ariel can't memorize one course. We keep a fixed showcase course and some layouts it has never seen for evaluation. We also trained a copy of the robot with every limb locked at normal length, on the same course with the same reward. That locked robot is the fixed-body rescuer from our inspiration, so we can see what reshaping actually buys.
The tools. Physics runs in MuJoCo through dm_control, training runs in PyTorch, and all of it ran on GPUs through Modal. For the showcase shots we exported the recorded simulation into Blender, so the nice-looking footage is the same motion the physics produced.
Challenges we ran into
Walking from scratch. Our first plan was to teach the humanoid to walk ourselves. It went nowhere, so we switched to the frozen MoCapAct walker and only trained the decision-making on top.
The catapult. With no limit on how fast a limb could extend, Ariel discovered it could launch itself by snapping its legs out. The speed cap fixed it.
Reward cheating. We assumed Ariel would cheat, so we blocked the usual tricks before training started: diving forward to grab progress, rocking back and forth to farm it, and vibrating along the ground to skate. Progress only counts at the pelvis, and only when it beats its own best. Every episode also logs how far the robot sank into obstacles, and clipping through one ends the run, so we could catch it passing through walls instead of going around them.
Crates. We expected the low deck to be the hard part. It turned out to be the crates. Late in training, both robots still got past them less than half the time.
A walker that only knows one body. The frozen walker learned on a normal human shape. It coped when one limb changed length, but it fell over once both leg segments went much shorter or longer than normal, so our policy had to learn to keep it balanced while it reshaped.
Moving the goalposts. We made the course harder twice while the robots were training (crates moved into the walking line, a steeper slope). Each run picked up from its last checkpoint instead of starting over, because there wasn't time to restart.
Accomplishments that we're proud of
Ariel learned to reshape itself with the training hints turned off. On our first course (low deck, gap, small rise), once the hints had faded to zero, the reshaping robot reached the astronaut in 305 of its last 400 training attempts (76%). The locked-body copy, trained the same way on the same course, reached it in 261 of 400 (65%). These are training attempts on randomized layouts under Earth gravity, not a separate test on unseen layouts.
On the harder course with crates in the path, both robots were still learning when we stopped, and full runs almost never succeeded yet. Per obstacle, though, the reshaping robot got under the low deck 91% of the time against 70% for the locked one, and up the slope 80% of the time against 55%.
We recorded every training attempt, so you can watch it learn instead of taking our word for it.
We scrapped our first robot at 1:14 AM, rebuilt on top of MoCapAct, and by morning had seven training runs going, the longest past 34 million simulation steps.
What we learned
Reward design took more of our time than the learning algorithm did. PPO mostly worked out of the box. Deciding what counts as progress, and blocking every way to fake it, is where the hours went.
Knowing what not to train from scratch saved the project. Our homemade walker learned to scramble and crawl and nothing else. Borrowing a walker that already worked let us spend the night on reshaping instead.
A baseline is what turns "it can reshape" into "reshaping helps." Without the locked robot, a 76% success rate tells you nothing.
The hard part isn't always where you think. We designed around the deck and the gap, and the crates are what held both robots back.
What's next for Ariel
Ariel only exists in simulation, and a simulated humanoid is a long way from a robot in a real collapse. People are working on that gap. NASA JPL has been testing EELS, a snake-like robot built to move through ice and tight spaces, and DuAxel, a rover that splits in two so one half can rappel down a cliff while the other anchors it.
Our next step is adding loose debris Ariel can push aside, and harder layouts it hasn't seen. Further out, we'd like to try the idea on real hardware with telescoping actuators.
We spent a lot of last night watching a robot learn to crouch under a deck. That's a long way from a building after an earthquake, but the people stuck in those gaps were on our minds the whole time.
How we used AI
We used generative AI (Claude, through Claude Code) to support our own work. We did the research, the design and most of the implementation ourselves, on top of open-source libraries.
During the build, AI helped us find and fix bugs, and when something wasn't working it suggested directions to try. It also helped us clean up a messy repo structure. Later in the night we set up AI agents to review the codebase and check that the RL training runs were behaving correctly.
After the build, we used it to lay out our presentation and to help edit our submission video. For this Devpost page, we described the project and our inspiration out loud using voice mode, and the AI drafted these answers (Markdown with LaTeX math) from what we said and from our codebase. We then edited them ourselves.
Credits
MoCapAct (Wagener et al., NeurIPS 2022; MIT code, CDLA-Permissive-2.0 models) · CMU Graphics Lab Motion Capture Database · MuJoCo and dm_control (Google DeepMind, Apache-2.0) · PyTorch · Stable-Baselines3 (MIT, used only to load MoCapAct's policy) · Blender · NASA Visible Earth imagery · compute by Modal. Inspired by Interstellar (dir. Christopher Nolan, 2014), the rescue robots Robin Murphy's team (later CRASAR) used at the World Trade Center in 2001, and NASA JPL's EELS and DuAxel robots.

Log in or sign up for Devpost to join the conversation.