Demo Video Link
https://drive.google.com/file/d/1IoqXAbmpUQyxEkxgziYTa03x_jNTlQdw/view
Inspiration
With the rise of code-generation tools like v0 and Bolt, anyone can generate a full-stack web application in seconds. However, we noticed a massive bottleneck: when multiple people or agents try to test, modify, and review changes on the same app, staging environments break, code gets overwritten, and security credentials get exposed. We were inspired to build a "detonation chamber" for web code, a place where humans and AI agents can build and test apps in isolated, secure, and parallel sandboxes.
How we built it
Sponsors Used: Daytona, ElevenLabs, Fireworks AI, Braintrust
SandStorm uses Daytona as its load-bearing infrastructure for secure, parallel QA. We deeply integrated the SDK to handle all orchestration, tunneling, and execution:
Agentic Browser Control: Let AI agents autonomously drive and verify complex UI workflows using virtual mouse and keyboard actions.
Millisecond Sandbox Forking: Cloned active environments instantly to fan out massive, concurrent QA test swarms.
Live Preview Tunneling: Securely exposed internal dev ports to route real-time app previews to external human testers.
Ephemeral QA Environments: Spun up perfectly isolated, customizable containers on the fly for individual testing sessions.
Isolated Test Execution: Safely orchestrated background compilations, dev servers, and Playwright test suites entirely within the container boundary.
Real-Time Code Injection: Pushed live code patches and extracted evaluation reports seamlessly via direct binary streams.
We built SandStorm using Next.js for the administrative console and dashboard, and Vite for the storefront checkout application under test. We integrated Daytona's SDK to spin up sandboxes and enforce egress firewall rules. We routed microphone audio through ElevenLabs Scribe to convert voice comments into coding instructions, which were fed into Fireworks AI's DeepSeek-v4-pro model to write filesystem patches. Finally, we set up Braintrust to run Axe accessibility audits and pixel-matching screenshot comparisons across parallel test runs.
Challenges we faced
The biggest challenge was handling resource management and API concurrency limits when running a "swarm" of tests. Spinning up 8–10 sandboxes simultaneously to run heavy Playwright and visual tests is CPU-intensive. We solved this by building a custom Resource-Safe Scheduler in the backend that throttles active sandboxes to a ceiling of 4 at a time, queueing the rest and re-running failures sequentially to isolate flaky runs.
Accomplishments that we're proud of
We are proud of building a fully-functioning end-to-end loop where you can literally speak a bug fix into existence, watch the code change on screen instantly, and see a parallel swarm of sandboxes validate it in seconds. We are also proud of our security containment design: using Daytona's Secrets Manager, we successfully mask production keys so that untrusted AI agents can test APIs without ever seeing or leaking the real credentials.
What we learned
We learned how powerful combining sandboxed environments with LLM agents can be. Before this hackathon, we viewed sandboxes merely as remote developer environments. Now, we see Daytona sandboxes as fundamental "security containers" and "detonation cells" that allow developers to execute unsafe agent loops, run parallel QA test runners, and isolate third-party integrations safely.
What's next for SandStorm
Next, we want to expand SandStorm to support larger-scale repositories, automate visual regression tests across multiple screen sizes and viewports, and expand our computer-use agent capabilities to handle complex user flows like multi-factor authentication (MFA) and CAPTCHA testing.
Built With
- braintrust
- daytona
- elevenlabs
- fireworksai
Log in or sign up for Devpost to join the conversation.