Inspiration
Physical AI breaks in ways a data center never does. The link drops, or the cloud goes slow instead of dying cleanly, or the device gets hot and starts skipping frames. When that happens, most systems fall back on one of two plans, and both worry me. Either someone wrote the backup by hand and never tested it, or an AI improvises in the moment on a machine that can hurt people. I wanted to try the opposite. Use AI ahead of time to find the backup plans, prove each one actually works, and put nothing on the device that has not passed.
What it does
Failsafe is a resilience compiler for edge and Physical-AI systems. You hand it an application and the promises it has to keep, for example, catch at least 95 percent of the people who enter a restricted zone and raise the alert within 2 seconds. From there it does five things. It proposes cheaper ways to run, like a lower resolution, fewer frames per second, a shorter cloud timeout, or dropping the cameras that do not matter. It breaks the world on purpose, cutting the network, slowing the cloud, choking the bandwidth, and loading the CPU. It measures each option in real time, at true wall-clock speed, because latency only means anything at 1x. A fixed checker, not the AI, decides which options actually keep the promises. Then it compiles the ones that pass into a small policy the device follows as a plain lookup, with no AI anywhere in the failure loop.
If a failure comes up that has no proven-safe option, the device does not guess. It drops to the most careful fallback and says NO VERIFIED MODE AVAILABLE.
The main finding surprised me, and it is all measured. A slow cloud is more dangerous than a dead one. With the same setup, a cloud that is fully offline gets handled by the fallback right away, so alerts still land on time at a p95 of 0.84 seconds. A cloud that is still reachable but slow is worse, because the pipeline waits out its timeout on every frame and the alert shows up too late, at a p95 of 5.05 seconds. A cloud with a 1.5 second round trip sits in between at 2.95 seconds, and still fails. The fix Failsafe finds is simple once you see it: shorten the cloud timeout so a slow cloud gets treated like a dead one, which pulls the slow case back under the 2 second limit.
How I built it
The whole loop runs from the command line before there is any interface. The test footage is a seeded synthetic corpus that drops real cut-out people onto drawn facility scenes, so I know the ground truth exactly, which means recall and precision are measured rather than guessed. Real video is kept separate as a holdout that the search never gets to see.
The workload itself is a real pipeline: a YOLOv8n person detector, foot-in-the-zone geometry, an optional cloud confirmation step, and an alert. It replays at true speed, and fault injectors sit around it to cut the network, slow the cloud, squeeze the bandwidth, and load the CPU.
NVIDIA Nemotron does the thinking that happens before deployment. It proposes the candidate modes and reads back the results through its OpenAI-compatible API, and it never gets to decide pass or fail. In one search it worked out the slow-is-not- dead idea on its own and offered the short-timeout fix as its very first candidate. I developed against NVIDIA Build, and the same model runs on Nebius Token Factory with a one-line change to the base URL.
A deterministic evaluator does the deciding. The compiler only accepts a mode when at least two clean runs across at least two different corpus seeds all pass, and it reports the worst of those runs rather than the best, so the numbers stay honest. For the cloud confirmation step I plugged in a real NVIDIA multimodal model, Nemotron-omni, which looks at the frame with the zone and the box drawn on and says whether the person is inside. Every answer is cached by a hash of the image, so replays stay identical, and the same code points at Cosmos Reason on a GPU node with one environment variable. Every result also records its seed, hashes, git commit, versions, and backend, so anyone can reproduce it or score it again later.
Challenges I ran into
The hardest one was making the word verified actually mean something. My first version signed off on a mode after a single run, which is really just a coin flip. I rebuilt it so a mode has to pass several clean runs across several seeds, and I made it report the worst run instead of the best. A couple of modes that had looked fine got thrown out, which was the point.
Measuring honestly on a laptop was harder than I expected. Real-time latency on a shared machine is fragile, and early campaigns kept getting polluted by background indexing, by the machine throttling after a heavy run, and by leftover CPU-load processes I had not cleaned up. So I added per-host calibration, a check that waits for the machine to be idle before each run, contention detection, a cooldown after the heavy runs, and a flat refusal to measure on a warm machine.
The last one was access. Nebius has no billing option for my country, Ghana, so Token Factory never activated and no paid Serverless Job has run. The integration is built and dry-run tested, but I want to be straight that it has not run for real yet. Cosmos Reason has the same shape of problem, since it needs a GPU node to self-host, so the confirmer runs on the hosted Nemotron model today and swaps to Cosmos the day a GPU is available.
Accomplishments that I'm proud of
I ran more than 200 real-time experiments across two full 72-configuration grids, with repeats over three corpus seeds. Out of that came a compiled, fail-closed policy with verified modes for the normal, slow, reachable-but-slow, dead, and low-bandwidth cloud cases, plus an honest refusal under heavy compute load, where nothing in the space passes and the system says so. The end-to-end run holds the mission through a full network outage, with overall recall of 0.982 and a p95 of 782 milliseconds, then recovers to normal once the link comes back. And the demo covers five domains, a warehouse camera, a self-driving dashcam, a drone, a robot's safety zone, and a robot arm, where four find a verified backup and the fifth correctly refuses.
What I learned
The hard part of resilience turned out not to be the model. It was being honest about the measurement. And the failure that actually bites is not the clean outage everyone tests for. It is the part that is slow but still alive.
What's next
Turn on Nebius Token Factory and run the parallel campaigns on Nebius Serverless Jobs once the regional billing is sorted, move the cloud confirmer from the hosted Nemotron model to Cosmos Reason on a GPU node, and keep growing both the synthetic corpus and the real-video holdout.
Log in or sign up for Devpost to join the conversation.