Inspiration

Every creative tool has a layer of paperwork nobody is excited to do. For SillyTavern creators, that layer is character-card schemas, lorebook triggers, injection positions, ordering, and PNG metadata.

We kept running into the same mismatch: the interesting work was inventing a person and a world, while the time-consuming work was packaging both into files that SillyTavern could reliably use. A card can load while still containing the wrong fields. A lorebook can be valid JSON but activate at the wrong time. A PNG can display correctly while its embedded character data is missing or damaged.

We built ST-Agent so the human keeps the taste, intent, and final approval, while the agent handles the repeatable work between an idea and a usable character package.

What it does

ST-Agent is a local Windows application for independent role-play creators, interactive-fiction writers, and narrative designers.

A creator chooses a local workspace, describes a story idea in English, and can optionally attach drafts, existing character cards, lorebooks, or a portrait. One agent asks a focused question only when a real decision is missing. Once the creator confirms the brief, it builds the requested package:

  • a SillyTavern-compatible Character Card V3 JSON file;
  • an optional PNG character card made from the creator's portrait;
  • a standalone Lorebook, an embedded Lorebook, or both;
  • persisted source, draft, and final-text files that remain readable outside the app.

The download section is built from verified files on disk, not from what the model claims it produced. If the required checks do not pass, ST-Agent does not say the case is complete.

How we built it

ST-Agent uses Python 3.12, the Strands Agents SDK, Streamlit, Pydantic, Pillow, and project-owned deterministic services. The interface and agent run in one local process, while model configuration remains in the sidebar through an OpenAI-compatible endpoint.

Our core architecture rule is:

The model decides what to do next; deterministic code decides what is true.

The Strands agent reasons about the creator's intent and selects among eight narrow tools: save_brief, save_content_document, build_character_card, build_lorebook, build_png_card, check_official_sources, validate_deliverables, and finish_case.

Those tools do not expose a shell, arbitrary filesystem access, or code execution. Each has typed inputs, one responsibility, and a structured result. Filesystem access is limited to the creator's selected case folder, uploads are identified from their bytes and size-limited, and network access is restricted to the configured model provider plus a small allowlist of official format references.

Accepted input is persisted before a model call begins. Case state lives in plain local files, important writes use atomic replacement, and each invocation reconstructs context from disk. Strands hook events power the live progress UI, but a sanitizer removes prompts, payloads, credentials, and hidden reasoning before anything is displayed.

At the end of the workflow, deterministic validators inspect the actual exported bytes. finish_case can succeed only when every requested artifact exists, has current lineage, passes the relevant JSON or PNG checks, and satisfies the delivery gate.

Challenges we ran into

The first challenge was drawing the authority boundary. A fully scripted wizard could not adapt to uneven creative input, but an unconstrained model could not be trusted with real files. We needed the agent to choose actions without letting its prose become completion authority.

The second challenge was format compatibility. Character Card V3, SillyTavern Lorebooks, and PNG metadata have related but distinct structures. We separated canonical creative content from deterministic serializers and added semantic read-back checks so a file is validated after packaging, not only before it.

The third challenge was recovery. Model calls can time out, Streamlit reruns the application, and a laptop can disappear halfway through a write. Treating files as memory, using atomic writes, and resuming from the earliest incomplete phase made interruption a recoverable state instead of a restart.

Accomplishments that we're proud of

  • A complete single-agent workflow with eight bounded, typed tools.
  • A deterministic delivery gate that rejects unsupported claims of completion.
  • Character JSON, PNG card, standalone Lorebook, and embedded Lorebook delivery.
  • File-based interruption and resume behavior without a database or hidden business state.
  • More than 200 automated tests covering units, formats, security boundaries, scripted-agent behavior, recovery, and the end-to-end controller path.
  • Windows CI, locked dependencies, a representative test case, architecture documentation, and a public MIT-licensed repository.
  • A public demo video and a three-part AWS Builder series documenting the build, delivery gate, and lessons learned.

What we learned

Narrow tools are easier for a model to select and easier for code to verify. Persisting accepted input before reasoning makes recovery dramatically simpler. Sanitized tool events are not just observability; they make autonomous work feel honest to the person waiting. A small confirm-before-write gate changes the relationship from “the agent acts at you” to “the agent works with you.”

Most importantly, agent tests need to exercise the loop, not only isolated functions. Scripted model responses running through the production controller let us verify question persistence, tool order, failure handling, delivery gating, and interrupted-case recovery without guessing from one long live run.

What's next for ST-Agent

This hackathon release is intentionally complete as a bounded local application; we are not attaching an imaginary cloud roadmap to a product that was designed around creator-owned files.

During judging, we will keep the repository, documentation, video, and published articles available. Afterward, real creator feedback will decide whether a future version should reduce setup friction or broaden its verified format coverage. Any expansion must preserve the boundary that made this release trustworthy: the model may drive, but only verified state may authorize delivery.

Build series

  1. We Built an Agent That Does the Boring Part of Interactive Fiction
  2. The Delivery Gate — Why Our Agent Isn't Allowed to Say “Done”
  3. Six Lessons From Building a Bounded Agent With the Strands Agents SDK

Built With

  • github
  • httpx
  • pillow
  • pydantic
  • pytest
  • python
  • strands-agents-sdk
  • streamlit
Share this project:

Updates

Submission history