Inspiration
It started with me (Adilet). A few of my API keys got exposed, and some were drained to $0, including ones I had loaded with $50 each. Then an AI coding agent deleted two projects I had built for myself. I hadn't been specific enough, it filled the gap with an assumption, and decided those projects should go. Nobody hacked me. I just gave an AI more access than the task needed and trusted it to guess right.
So I started digging, and it turned out I wasn't unlucky. I was early.
Last July, a founder building an app with an AI coding agent told it not to change anything. On day nine it deleted his live production database, then told him the damage couldn't be undone. A month later, attackers hijacked the popular Nx build package and, in the first documented case of its kind, used developers' own AI command-line tools to search their machines for secrets. Researchers counted 2,349 stolen credentials from 1,079 developer systems, and 85% of those machines were Macs.
Meanwhile, 84% of developers use AI tools, and only 33% trust the accuracy of what those tools produce (Stack Overflow Developer Survey 2025).
We build with AI agents every day, and so does almost everyone at this hackathon. Students, interns, freelancers and people who have never written a line of code now hand an AI access to the same laptop that holds their IDs, tax forms, passwords and family photos. That made us ask a simple question: what's the point of building something to solve a problem if building it creates a bigger one ?
What it does
Tini Family is a desktop app that lets you build with an AI agent without handing it your whole computer. It's a small 3D world with three characters:
- Tini reads your request and builds a fence around only the folders the job needs. You see it as one contract card instead of forty permission pop-ups.
- The Agent is Claude Code/Codex/Antigravity etc, building your project inside the fence.
- Tina inspects everything before it comes in and everything the Agent makes before it goes out. She's allergic to red: nothing launches until every fence segment is green.
The fence has three layers:
- Copies, not originals. Tini copies only the approved folders into a separate project folder, stripping GPS locations from photos on the way. Your originals are never touched. (Based on: Website built example)
- Every step checked. Every action the Agent takes, whether opening a file, saving one or running a command, is checked before it happens. Anything outside the fence is blocked with a plain-English reason.
- The Mac's own lock. Commands run inside the macOS kernel sandbox, which stops disguised tricks the step check can't see.
When the Agent needs more, it asks through a request door. Tina inspects the folder first. In our demo she checks 1,212 files, finds 1,199 personal photos and a driver's license, and recommends letting in only the 12 job-site photos. After each turn she scans the result for leaked API keys, ID documents, personal data and hidden photo details. Every red spot comes with a one-click fix, and a final access report shows everything that was allowed, blocked, narrowed and fixed.
It's built for multi-turn work: you keep asking for changes, and Tini only interrupts you when access changes.
How we built it
- The engine runs locally on your Mac in Node.js and TypeScript, and streams everything live to the app.
- The Agent is Claude Code, run through the Claude Agent SDK. Before every single action, our guard checks where it points, and the macOS kernel sandbox locks down every command it runs.
- Tini's planning is plain code that decides access. Gemini only helps with labels and suggestions, and code checks everything it says. If Gemini fails, Claude takes over automatically.
- Tina's inspections run on your own machine: gitleaks for leaked keys, ExifTool for hidden photo data, and PDF reading for ID documents. Nothing is uploaded to be checked.
- The app is a 3D world built with React and Three.js. The landing page runs on Vercel at our GoDaddy Registry domain, tini.cooking.
Challenges we ran into
- Our first security test lied to us. The very first test of the fence came back "passed." It had passed because nothing ran: our API key was rejected on every call, and failures were being counted as blocks. We caught it, made the test stricter, and ran it again for real.
- We rebuilt the product overnight. Halfway through Saturday we realized we'd designed for one prompt and one website. Real people keep asking for changes, on any task. We re-planned the whole engine around a multi-turn workspace while the clock kept running.
- The AI broke our demo by being too smart. Our planned demo depended on the AI falling for a hidden attack. Claude spotted it and refused on its own. Great for users, a disaster for our story. We rebuilt the demo honestly: the attack is now a clearly labeled simulation through the real fence, and the Agent got a proper way to ask for access instead of pushing.
- Then the fence made the AI too careful. Once it knew it was fenced, Claude started treating the client's own notes as suspicious. Our first real runs had nothing for Tina to catch, which meant no demo. We had to learn how safety changes an AI's behavior, and redesign around it.
- Our own safety tools fought us. The permission systems on our development tools blocked us from copying our own keys and running our own tests, and at one point flagged our fake demo bait as a real attack. We had to prove to our own tools that we were the good guys.
- Gemini went down on us, again and again. Quota errors, overloaded servers and requests that hung silently for minutes. So we built a fallback ladder that never gives up.
- We fell six hours behind. On Saturday evening we cut five features in one decision and focused on what the demo truly needed.
- And we barely slept. About six hours total across 36 hours.
Accomplishments that we're proud of
- The fence holds, and we proved it. We tried to sneak past our own step check with a disguised command, and the Mac's sandbox still stopped it.
- Even a tricked AI can't open the gate. We fed our planner a malicious answer asking for SSH keys and extra folders. Our code threw out every single request.
- Tina catches real leaks. In live runs the Agent put a client's API key into a website and published a photo carrying the owner's name. Tina caught both before launch, every time.
- It never stops working. If Gemini goes down, Claude takes over. If both go down, it still answers. 100 out of 100 in our stress test.
- Faster without being weaker. We cut the first build from about 3 minutes to under 2 without removing a single safety check.
- Two people, one weekend: the engine, a 3D app, a replay system and a landing page.
What we learned
Most of AI safety is plumbing, not magic. The hard part is making product automatic and understandable for someone who has never opened a terminal. We also learned to keep AI out of every decision that matters: AI is great at labels and explanations, and code should decide what gets in.
What's next for Tini Family
- Your existing projects. Today the Agent works on copies in a new folder. Next, it works right inside the projects you already have, still fenced.
- Asking for websites, not just folders. If the Agent needs to reach a new website, you'll get the same simple "allow or deny" card.
- Tina learns to see. With your permission, she'll check the pictures themselves for faces, license plates and scanned IDs, not just their hidden data.
- Windows and Linux, not just Mac.
- Teams. A company sets one fence for everyone and gets a report of what every AI touched.
Built With
- anthropic
- claude-agent-sdk
- claude-code
- exiftool
- express.js
- framer-motion
- gemini
- github
- gitleaks
- godaddy-registry
- google-ai-studio
- google-cloud
- google-genai-sdk
- macos
- node.js
- pdf.js
- phaser.js
- react
- react-three-fiber
- socket.io
- three.js
- typescript
- vercel
- vite
- zod
Log in or sign up for Devpost to join the conversation.