Inspiration
What it does
How we built it
Challenges we ran into
Accomplishments that we're proud of
What we learned
What's next for Cavernum: Number Puzzle
Inspiration
I've been playing games for 30 years and never built one. I had an idea for a number-pairing puzzle mechanic, kept sketching levels in my head, and eventually decided to just make it instead of thinking about it. I don't code. That wasn't a reason to stop, just a constraint on how I'd have to build it.
What it does
Cavernum is a cave-themed number puzzle. You clear tiles by pairing two that are equal or add up to ten, as long as a clear path connects them. Sixty levels, and it doesn't stay easy: rocks, gates, and locked chests get added as you go, and the boards are hand-placed rather than randomly generated, so the difficulty is deliberate instead of accidental.
How I built it
ChatGPT built the visual world. Claude worked out the mechanics and level rules with me before any code existed, then turned that spec into scoped, bounded prompts. Replit Agent wrote every line of the actual code. My job was running the loop between them: check what Agent proposed against the spec, catch what didn't match, send it back.
I never wrote code myself, but I ran the entire product and QA process by hand. I played the same handful of levels enough times that my wife started checking on me.
Challenges I faced
Procedurally generated levels were solvable but boring; matches sat right next to each other and blockers landed in uninteresting spots. I had Claude build me a level editor and switched to hand-placing every tile. That fixed the game and created a new problem: I couldn't get Claude to stop flagging bugs in the old generator code that nothing uses anymore. I explained that it was dead code more times than I explained the actual mechanics.
The native build had its own fights, mainly a Firebase/Firestore dependency conflict that broke the iOS archive step repeatedly before we landed on a stable configuration and stopped touching it.
The harder problem was trusting AI output at all. An agent once reported 231 out of 231 checks passed. True, and also decorative; the checks didn't test anything real. Nothing caught it automatically. I did, because the number looked too clean. After that I split execution from review, one agent does the work, a separate one checks the claims, and every incident like that became a standing rule.
Even with that in place, a merged and reviewed fix once almost shipped a worse bug than the one it patched, and it reached a real player before we caught it. Verifying that an agent did what it claimed is a different skill from verifying it was solving the right problem. The game is small enough that I know how it's supposed to behave, so a wrong-problem answer usually looks wrong to me on sight. That's not a safeguard I could hand off, and it's the main thing I'd tell anyone else trying this: the review discipline matters more than the prompting.
Built With
- admob
- android
- chatgpt
- claude
- expo.io
- firebase
- ios
- react-native
- replit
- revenuecat
- typescript
Log in or sign up for Devpost to join the conversation.