Our Inspiration

visionOS lets you build interactions that couldn't exist anywhere else. Not a screen you look at, but something you can touch and interact with in 3D. We started from two feelings that have nothing to do with computers: the moment just before a mystery box opens, and the plain physical joy of breaking something. Then we realised they're the same feeling cut in half.

What it does

Your hand is a water gun, and you blast a massive block of rock apart to uncover a 3D object generated inside it, one you described or one that surprises you. You can describe something and an on device AI brings it to life.

How we built it

Almost every image-to-3D system finishes by giving a mesh as an output. But TRELLIS.2 doesn't start with a mesh. Its first stage produces a sparse voxel grid, a cloud of solid little cubes, and only converts that into a skin at the very end. So we stop before the conversion and take the cubes. How the on device generative-AI works: your photo, then Vision cuts the subject out onto a black card, then DINOv3 boils it down to its essence, then a flow model starts with a cube of static and nudges it twelve times toward your object (Core ML where the device allows it, MLX behind it), then a decoder answers 262,144 yes-or-no questions, one per cube: solid, or empty air? Then a colour pass paints every surviving cube. Those cubes go straight into the grid you spray.

Challenges we ran into

Five gigabytes of RAM. VisionOS kills an app that goes much past 5 GB of memory, and a generative 3D model takes many times more than that. Nearly every hard problem this weekend was a consequence of that one number. We solved that by quantizing TRELLIS 2. Making it fit inside and runnable by Vision Pro.

Built With

Share this project:

Updates

Submission history