Inspiration
We are both game designers. More and more we look towards turn based games, not just as interesting surfaces to design on but as powerful tools to model the real world.
The initial idea was the make a tool to partially automate testing of game ideas, though what we ended up with was stronger than expected.
The implementation, which is explored more below, is a mad mashup of techniques from AI Automation, HCI and Control Theory, and allows for both exploration of existing games and prototyping novel ideas.
What it does
Given a description of a turn based game and its rules from the user, a Rules agent attempts to formalize the game into a prototype Python implementation, then deploys a swarm of player agents to play the game while giving commentary on their strategy. Player agents are aware of the prose of the game rules, the current state of their game instance, and their current legal moves, allowing them to amass insights on what they are doing and why. After all games are concluded, or a max number of turns reached, a final Design agent collects the insights made by the players and synthesizes a review of the design to return to the user.
The user can then modify their game and repeat the loop, or start fresh.
The goal is to automate the early testing phase of a game. But this also demonstrates some properties of a more general use model. Of course, the components are each LLMs, but their roles are important: formalizing the input (the "environment"), acting out ("predicting") how things will proceed, and reflecting on why things went the way they did are all considered in this loop. If you can phrase your problem as a game, this system can work on it. Though this particular system is more interested in whether a problem has artistic value or not, rather than finding a solution.
How we built it
We used Python with google's Gemini API, and lots of markdown.
To get around the challenges enumerated below, we set up an architecture where an initial rule agent formalizes the rules, then processes them further into deterministic python code. This code is then copied to each of the player agents as a source of truth, an internal model of the world for them to check back on for any questions about the game state, or what legal moves they have.
The second most important architectural detail is questions, the initial step can be interrupted if the rules agent believes there is something off about the input. Ambiguity or contradictions can lead to unexpected results, and in lists of rules modern LLMs are good at spotting these holes. If spotted, the user will get a ping with a question on how to resolve the discrepancy.
Challenges we ran into
The initial flow of execution was to simply describe a game to a chatbot, have it deploy agents to play the game and follow the rules, then summarize the interaction and give feedback on the design. However, LLMs are not good at this task for a few reasons.
One is spatial reasoning. Most games have a notion of space. However, even modern LLMs, with their hard coded foundations and extensive training in abstract math, struggle to keep up with concepts of distance and position, which most turn-based games rely on, in part or in full, as an intuitive store of information for the player.
More pressingly, chatbots don't play by the rules. Swarms are famously hard to coordinate. To get a useful result out of them, none of them can misalign from the original input ruleset.
Accomplishments that we're proud of
Much of the code in this project is unusually well tested for a hackathon project, with multiple iterations of development and high unit test coverage. Since we want to keep using this tool in the future, it's a little more well built than your average hack.
What we learned
We ran into a lot of the issues that arise when including an agent swarm in your architecture, and gained some valuable experience in solving those problems. We also learned about some of the security vulnerabilities that come with allowing a system to run arbitrary code.
What's next for BGent
We're both going to use this tool for our own purposes in design. It's hard to know if a game "works" before taking it out in the real world, and often the process of describing the rules fully is illuminating in its own right.
Log in or sign up for Devpost to join the conversation.