Inspiration

A game update can launch without a crash while silently corrupting the meaning of an old save. In the representative Orb Forge incident, a returning player starts with 425 coins, one completed Orb, and a craft at polish 2/3. Version 2.0 incorrectly shows an already-claimed Founder Gift. One legitimate click replaces that history with 200 coins, no completed Orb, and no active craft.

Crash reporting sees a successful launch. Screenshot comparison sees a valid screen. The developer still has to reproduce the historical state, determine what was lost, isolate the smallest valid save, find the responsible loader branch, limit the repair, and rerun the exact input. StateWeaver automates that repetitive, judgment-heavy pre-release investigation for indie and small game teams.

What it does

The local StateWeaver path loads the historical save into the public Orb Forge fixture in Chromium and captures DOM state, persisted JSON, screenshots, and SHA-256 evidence. StateWeaver uses a Strands agent in Amazon Bedrock AgentCore to investigate captured browser evidence, trace the regression, enforce human approval, and verify recovery with the identical historical save.

The deployed agent compares player-owned meaning, minimizes the counterexample, investigates the relevant new-field semantics, traces the source location, and creates a bounded repair preview. It does not launch arbitrary uploaded games or arbitrary candidate builds.

The model cannot approve its own change. Before human approval, the repair endpoint returns HTTP 409, changed=false, and leaves source unchanged. After approval, deterministic policy permits only game.js. StateWeaver then rechecks the byte-identical save and reports success only when 425 coins, one completed Orb, polish 2/3, and duplicate Founder Gift protection are all preserved.

How we built it

The official Strands Agents SDK runs an evidence-driven investigation loop using Amazon Nova Pro in Amazon Bedrock AgentCore Runtime. The model chooses what to inspect next and records why. The prompt supplies goals, available capabilities, safety rules, and completion conditions—not a fixed tool order.

Deterministic code separately controls the save schema, semantic invariants, approval token, allowed file scope, SHA-256 identity, idempotency, and final verdict. The bounded anonymous gateway uses AWS Lambda and DynamoDB. Locally, Playwright and Chromium execute the Orb Forge fixture and capture the actual browser evidence.

Why it is different

StateWeaver is not general game QA and does not merely report a bug. It targets silent semantic save regressions: cases where the game looks healthy while a returning player's accumulated history no longer means the same thing.

The proof uses the same SHA-256 input that failed before repair. StateWeaver does not substitute a newly generated save to make the result pass. The agent reasons; deterministic guards decide whether a change is allowed and whether recovery is truly complete.

Challenges

The hardest boundary was separating model judgment from safety authority. We redesigned the investigation so the model selects tools from evidence instead of following a scripted order, while deterministic validation rejects missing evidence and every unapproved or out-of-scope change. We also constrained the public demo to one fixture, one concurrent investigation, 30 calls per UTC day, a one-use approval, and game.js as the only repair target.

Evidence and limitations

  • 5/5 fresh anonymous API round trips
  • 30/30 defects detected
  • 0/12 healthy-save false positives
  • 42 synthetic cases across six frozen regression classes
  • 0 unauthorized changes

The benchmark is synthetic and deliberately limited to six frozen regression classes. The AgentCore deployment is verified on the representative incident; this is not evidence of arbitrary-game generalization or measured time savings. No external game developer validation has been completed.

Links

Built With

Share this project:

Updates

Submission history