Inspiration
Mars is between 3 and 22 minutes away from Earth at the speed of light, each way. That means every time a rover hits something unexpected, like a blocked path or a rock too hard to drill, it stops and waits for Earth to respond. A single question can cost a full round trip, and missions lose days this way.
You can't make the signal faster. But you can change how often a robot has to wait for it. The hard part isn't building autonomy; it's deciding how much autonomy a specific robot can be trusted with at a specific site and time. Today that decision comes from engineering judgment and review meetings. I wanted to see if it could come from data.
What it does
Leeway figures out how much a Mars robot can safely decide on its own, using real NASA data, and then runs that decision in a live mission control.
1. The autonomy envelope agent. You describe a mission, for example "Rover at the Jezero delta front, 30 sols." The agent loads real data for that site and time window: HiRISE terrain of Jezero crater, JPL Horizons ephemerides for Earth–Mars light-time and communication windows, and the sun's position on Mars. It proposes an autonomy envelope (which situations the rover may handle alone, its safety limits, and when it must escalate), stress-tests it across [2,000] simulated sols, and tightens it wherever a run ends badly. It stops at the widest envelope with zero unsafe outcomes, then re-checks on fresh runs it was never tuned on. The output is the envelope itself plus an evidence report showing why each limit exists.
2. Mission control. The envelope loads into a live mission control that runs it side by side with conventional operations on the same terrain, clock and surprises. When the path is blocked, the conventional rover stops and waits a full round trip. The Leeway rover already has a pre-approved branch and keeps driving. When it meets something truly outside its envelope, it stops safely and sends one escalation designed to be answered in a single reply: what happened, its options, and its recommendation.
3. Bandwidth-light situational awareness. Instead of downlinking a full image, the rover sends a few hundred bytes describing the scene. On the ground, Grok Imagine reconstructs it so the operator can see what the rover sees. Reconstructions are always labeled as illustrative, and decisions run only on the structured data.
In simulation, the same mission went from [X] round trips under conventional operations to [Y] with Leeway.
How we built it
- Frontend: React + TypeScript (Vite), with three pages: a landing page with a live Earth–Mars light-time line, mission control, and the envelope agent page.
- Backend: Node/Express, which runs the simulation workers and keeps all API keys server-side.
- Simulation engine: a deterministic onboard executor that runs plans with branch conditions and escalation triggers, a delay-injecting message link between Earth and Mars, and a headless fast-forward mode for running thousands of campaigns.
- Plan schema: one strict, validated (zod) format shared by the plan compiler, the safety validator, the executor and the UI.
- Grok: compiles plain-language mission intent into structured contingency plans, and acts as the agent that proposes, tests and tunes autonomy envelopes using the simulator as a tool.
- Grok Imagine: reconstructs scenes on the ground from the rover's compact text descriptions.
- SpacetimeDB: the real-time data layer. Real data sources with provenance, per-sol conditions, the agent's live log, campaign results, failures and versioned envelopes all live there, and the pages update through subscriptions. Loading an envelope into mission control is a single reducer call.
- Real data: HiRISE digital terrain models of Jezero crater (processed into heightmaps with slopes and hazards), JPL Horizons ephemerides, and NASA's Mars24 algorithm for sun position and season.
- Cursor for the entire build.
Challenges we ran into
- Framing the problem honestly. "Reducing latency" isn't possible, so the project had to be about reducing round trips and time lost to the delay. That framing shaped every design decision.
- A fair baseline. Real rovers aren't joysticked step by step, so the conventional baseline runs a full command sequence with no contingencies and stops on any surprise. A weaker baseline would have made the results look better and meant less.
- Making the agent trustworthy. An agent that tunes its own safety limits can overfit to its own test runs. I added a final check on fresh random seeds and strict iteration and run budgets.
- Visualization. I tried a 3D first-person Mars view and a live pipeline that analyzed Perseverance's latest raw images. Neither met the bar for reliability in time, so I cut them rather than show something shaky.
- Learning SpacetimeDB mid-hackathon, especially reading results through subscriptions instead of return values.
Accomplishments that we're proud of
- Real NASA data drives actual decisions: change the site or the date, and the terrain limits, communication windows and resulting envelope all change.
- An agent that does real work in a loop: propose, simulate, inspect failures, revise, and verify.
- Escalations designed around the cost of a round trip, so a single reply is always enough.
- Building the whole system solo, with a clear line between what's real (the software and the data) and what's simulated (the mission).
What we learned
- The best use of AI here wasn't a smarter robot. It was a better answer to "how much should we trust the robot?"
- A shared, strict schema early on saved hours later, because every component spoke the same format.
- Honesty is a feature with technical judges: stating that results are simulated, labeling AI reconstructions, and calling the envelope a recommendation rather than a certification all made the project more credible.
- Real-time shared state is a good fit for agent work, because people can watch the reasoning happen instead of waiting for a final answer.
What's next for Leeway
- Earned autonomy: learn from how operators answer escalations, and propose evidence-backed envelope expansions that a human approves.
- Benchmarking against real missions: replay stretches of Perseverance's actual traverse to compare conventional operations with Leeway's model.
- More sites and higher-fidelity simulation: additional HiRISE sites, better terrain physics, and dust-season effects.
- Live scenes: revisit planning against Perseverance's latest raw images, with stronger vision and a reliable pipeline.
- Beyond Mars: the same approach applies anywhere robots operate with long communication delays, such as the Moon's far side, outer-planet missions and deep-sea work.
Built With
- cursor
- express.js
- grok
- hirise
- imagine
- javascript
- nasa
- node.js
- python
- react
- spacetimedb
- typescript
- vite
- xai
- zod

Log in or sign up for Devpost to join the conversation.