Inspiration

Inspired by the “Fetching Reality” challenge, we wanted to turn a static photo of any living creature into a companion you could spend time with using an ordinary webcam and a projected 3d space.

What it does

Users can generate and rig a model from a photo, inspect its skeleton, and explore animated characters in first person.

Your real hands become virtual hands. Wave to get a response, wave a character to come closer, or reach out to pet it. grab a ball, throw it, and watch your companion chase it, pick it up, and bring it back.

The experience includes a dog, Pikachu, Lebron James, and Keanu Reaves, each with different reactions. Your companion lives in a 3D environment generated from a single photograph, giving you an outdoor space to explore, play fetch, and spend time together.

How we built it

We used TypeScript and Three.js for rendering, animation, and interaction, and Next.js/React for model generation and the viewer.

Our character generation pipeline uses an image-to-3D model hosted on fal.ai, followed by Tripo’s rig compatibility check and automatic rigging. GPT 6 astra/sol then creates animations based on the rigged 3d model and also defines the characters' behaviour.

MediaPipe tracks both hands through the webcam. We smooth the landmarks, stabilize gesture recognition, and estimate reach from changes in apparent hand size.

We used World Labs’ Marble 1.1 to turn a single backyard photograph into a 3D environment. The generation produced Gaussian splat representations. We used its textured mesh export to bring the backyard into Three.js. We aligned the environment with the character and interaction systems so the pet, virtual hands, and ball share the same space.

We also used Python with NumPy and Pillow to process the generated backyard’s collision data, build navigation maps, and render diagnostic previews for visual verification.

Challenges we ran into

Mapping webcam input into a 3D world was one of our biggest challenges. A hand can overlap a character on screen while still being too far away to touch it. We had to align mirrored coordinates, virtual perspective, estimated depth, and contact detection so interactions felt consistent.

It was difficult to get GPT to produce reliable animations and behaviour since different rigs had different bone names, proportions, and facing directions. We needed believable reactions and smooth transitions while keeping feet planted and avoiding accumulated deformation.

Fetch connected nearly every system. Throws needed to reflect hand movement without gaining extra force from camera movement. Characters needed to retrieve the ball and offer it back even when the user interrupted the sequence or hand tracking disappeared.

Accomplishments that we're proud of

finishing on time.

What we learned

Making a character feel alive takes more than a good model. Timing, contact, motion transitions, and small idle behaviors all shape how users perceive it.

What's next for Pet Maker

We want to make the path from uploading your own pet photo to playing with it more seamless and faster.

Built With

Share this project:

Updates

Submission history