Tracks: Bring Your Own Project The Code Registry Challenge

UNACCOUNTED is being submitted to the Bring Your Own Project track as the primary hackathon project, and also to The Code Registry Challenge for code quality, dependency, and security evaluation.

Inspiration

I came into this hackathon wanting to learn by building something I would actually be excited to keep working on afterward.

I love science fiction, mystery games, and systems where the player has to piece together what happened rather than simply follow a scripted sequence. At the same time, I wanted to experiment with AI agents in a way where the AI was doing something meaningful, not just adding a chatbot to an existing application.

That led to UNACCOUNTED, an AI-agent-driven science-fiction investigation game.

The player is the Mission Commander of a classified investigation team sent to investigate abandoned ships, stations, and other deep-space environments.

For this hackathon I built the first playable investigation, Case 037: SNV Petrel.

The Petrel is a government survey vessel found drifting in deep space. Auxiliary power is still running, no life signs are detected aboard, and several lifeboats are missing. The player's job is to send specialists to investigate the ship, recover evidence, reconstruct what happened, and submit a final evidence-backed finding.

The central design question became:

How can I let a player interact naturally with AI agents without allowing the language model to rewrite the mystery?

That became the core architectural principle of the project:

AI interprets player intent. The deterministic game engine owns truth.

What I Built

UNACCOUNTED is a complete playable vertical slice that moves through:

classified briefing → Mission Control → investigation → evidence recovery → final finding → deterministic evaluation

The player can issue conventional investigation commands or use natural language.

For example:

Engineering, inspect the reactor while Science reconstructs the accident.

A locally running language model interprets that command and decomposes it into bounded specialist tasks.

In this example, it can assign:

  • Engineering → inspect the reactor
  • Science → analyze environmental and navigation records

Those actions are then validated by the backend before the deterministic game engine executes them.

The important distinction is that the model does not generate the evidence.

Evidence such as reactor shutdown records, thermal-system analysis, navigation data, and lifeboat telemetry already exists as canonical case data. The AI helps the player decide what specialists should investigate, but the underlying facts never change based on what the model says.

Once enough evidence has been recovered, the player can submit a formal incident finding. They must choose a conclusion and cite the evidence supporting it.

That evaluation contains no AI at all.

A correct guess without the necessary supporting evidence is not considered a supported finding.

How I Built It

The frontend is built with:

  • Next.js
  • React
  • TypeScript
  • Tailwind CSS

The backend uses:

  • Python
  • FastAPI
  • Pydantic
  • JSON-based case data

For the AI layer I used:

  • Qwen3 4B Instruct
  • Ollama

The model runs completely locally on my Mac. There is no paid API call or API key required during gameplay.

The model receives a deliberately limited set of specialist roles and permitted actions. Its output is constrained to a structured schema and validated with Pydantic.

The backend then independently verifies that every requested role/action combination is allowed before any task executes.

For example, Engineering can inspect the reactor, but Security cannot.

If a model response contains one valid task followed by an invalid task, the entire command is rejected before either task can change mission state.

The same deterministic investigation engine powers both natural-language commands and manual controls.

Challenges

The biggest challenge was deciding where AI should stop.

It would have been much easier to give the model the entire case and ask it to generate discoveries dynamically. But for a mystery game, that creates a serious design problem: the model could hallucinate evidence, contradict earlier facts, or effectively change the answer while the player is investigating it.

Instead, I treated the language model as a constrained semantic interface.

That required separating three concepts:

Truth → Evidence → Interpretation

Truth is what actually happened.

Evidence is what the player can discover about what happened.

Interpretation is what the player or AI agents believe that evidence means.

Keeping those separate made the investigation far more reliable.

Another challenge was running the model locally. I ultimately used Ollama with Qwen3 4B, which proved to be a good fit because the model is solving a narrow routing problem rather than generating the entire game.

I also had to keep the hackathon scope under control. There are many directions I want to take UNACCOUNTED, but I focused on completing one investigation end to end instead of building many partially finished systems.

What I Learned

The biggest lesson was that building an agentic system does not necessarily mean giving an AI unlimited autonomy.

In this project, the interesting behavior comes from allowing the player to delegate goals naturally while keeping the agents inside explicit capabilities and permissions.

I also learned how useful structured model output can be. Combining a language model with schema validation and deterministic server-side rules made the AI component much easier to reason about and test.

Before finishing, I added a backend reliability suite with 26 automated tests covering game-state behavior and the AI boundary, including malformed model output, invented actions, invalid role/action combinations, unrecovered evidence, incorrect conclusions, reset behavior, and multi-agent commands.

All 26 tests pass.

What's Next

This hackathon build represents one investigation, but the architecture is designed to support a much larger game.

Future cases could add:

  • persistent specialist characters
  • character memory and shared mission knowledge
  • equipment and capabilities
  • larger ships, stations, habitats, and planetary facilities
  • hazards and consequential decisions
  • concurrent specialist tasks
  • incomplete or corrupted evidence
  • multiple investigations with different underlying causes

The long-term goal is to create a science-fiction investigation game where AI makes commanding a team feel natural and unpredictable, while deterministic ground truth keeps every mystery fair.

The AI helps you investigate. It does not get to rewrite the mystery.

Built With

Share this project:

Updates

Submission history