Inspiration

I ship a small fortune-reading app on Google Play alone. Every release repeats the same chores that are easy to describe and easy to get wrong: does the store listing still carry the AI-disclosure and "not advice" framing lines it needs? Which RevenueCat webhook event is this, and do I revoke the entitlement? Turn a 39-commit git log into release notes and a "what's new" block under Google Play's 500-character cap. Write the Korean listing I have been putting off because I cannot proofread it well.

None of these is hard. Each one costs an evening of judgment work, and skipping one gets a release rejected or a chargeback. I wanted an agent that does the work and then hands me the short list of things only I can decide — not one I have to trust blindly.

What it does

ReleaseKeeper is one Strands agent with four deterministic tools, run from a CLI or a one-page web UI on real release inputs:

Tool Input Output
disclosure_check store / result-screen copy red · yellow · green findings with a concrete fix each; verdict BLOCK / SHIP_WITH_FIXES / SHIP
refund_triage a RevenueCat webhook JSON body event class, what to do to the entitlement, and a ready-to-edit reply (en / ko)
release_notes git log --oneline New / Improved / Fixed notes plus a store "what's new" block under the Play cap
listing_brief → listing_check app facts JSON + target language a full store listing written by the model inside caps and mandatory disclosure lines, then checked deterministically

Every tool also runs without any model (--raw), so the deterministic checks are usable on their own and testable without a key. Findings are risk flags with a fix, never legal verdicts; anything not decidable from the inputs comes back under needs_input.

The design rule that runs through all four: the generative step is bracketed. A deterministic tool sets the constraints before the model writes, a deterministic tool measures what it wrote, and the CLI / web page re-run that measurement themselves and exit 1 on a red finding. The model plans, calls and explains. It never gets to be the last word on a rule.

How we built it

  • Strands Agents SDK: a single Agent, five @tool functions (plain Python, docstring = the contract the model reads), a per-tool system prompt. Each request is a single turn (prompt → tool → answer) because chat-completions endpoints do not carry reasoning content across turns; for release chores that is the honest shape anyway — there is a document and a verdict, not a conversation.
  • Model: LiteLLMModel pointed at an OpenAI-compatible endpoint. Default is an OpenRouter free-tier model, so the whole run costs $0; switching to Amazon Bedrock or any other endpoint is three environment variables (RK_MODEL, RK_API_BASE, RK_API_KEY), no code change.
  • Surface: one FastAPI process serves the web page (four tabs × "Run tool" / "Ask the agent"), the OpenAPI docs, and the Amazon Bedrock AgentCore Runtime contract (GET /ping, POST /invocations, port 8080, ARM64 Dockerfile). The AgentCore contract is implemented and verified locally; deployment was intentionally skipped for this submission (optional under the rules).
  • Tests: 51 unit tests, none of which need a model — because everything that decides anything is deterministic.
  • Demo video: reproducible from the repo (demo/): Playwright drives the real web page on the real inputs, ffmpeg assembles the clips; narration is synthetic text-to-speech and is disclosed as such.

Architecture diagram: docs/architecture.png in the repo (also attached to this submission).

Challenges we ran into

Everything below came from running the tools on my own live release (Saju Today 1.1.17), not from fixtures. The full list is in docs/e2e_saju_1.1.17.md.

  • The model does not count. Asked for a store block under 500 characters it returned 622 with a confident "(under 500 chars)" label. The fix was not a prompt: the CLI and page now measure the last fenced block themselves, fail with exit 1, and offer a version trimmed to whole bullets. On a later run the model skipped the code fence entirely, so the extractor learned a fallback — a checker that only parses the model's tidiest output silently passes on a bad day.
  • The model copies drafts. Handing it a field called store_whats_new got my raw commit subjects pasted into the store text. Renaming the fields draft_* and saying "developer jargon — rewrite, never copy" in the docstring changed behaviour more than any system-prompt edit.
  • The model does not translate what it can copy. An English refund line landed verbatim inside the Korean listing twice. The brief now substitutes the language's default line, and listing_check reds any English sentence in a Korean description.
  • RevenueCat has no REFUND event. A refund arrives as CANCELLATION with cancel_reason = CUSTOMER_SUPPORT; a handler written for a REFUND type never fires. The tool also warns that no refund event reaches you at all until the Play service account is registered in RevenueCat — code being right and detection being wired are two different checks.
  • A per-session commit history has no feat: / fix: commits. 38 of 39 commits in the real range were chore(session); the first run produced an empty note without saying why. The classifier now has a keyword pass and names the flag to use.
  • Free-tier reliability. Recording the demo hit "Upstream error … Service temporarily overloaded" (HTTP 502) four times; the recorder retries with a pause. That is the real price of the zero-cost path.

Accomplishments that we're proud of

  • Tested on a real release, with the numbers to show for it (one pass, default free model, wall-clock):
Step Real input Result Time
check live Google Play listing (3836 chars, Play Developer API) SHIP_WITH_FIXES — one real finding: no minimum-age line 0.0 s raw · 7.8 s agent
notes the real 39-commit release range 622/500 caught by the post-check → 364/500 on the next run ~50 s
listing --lang ko app facts from the live listing (the app had no Korean listing) full Korean listing, every field under cap, check re-run by the CLI agrees 22–71 s
refund RevenueCat CANCELLATION / CUSTOMER_SUPPORT payload (anonymized) class refund, revoke entitlement, reply drafted, Play-credentials warning 3.9 s
  • A full four-tool pass in the demo recording used about 32k input / 11k output tokens in total at $0 model cost.
  • 51 unit tests without a model. Public MIT repo with commits dated inside the submission period. Reused rule logic from my own pre-launch checklist is re-implemented and disclosed in the README.
  • The disclosure tool discloses itself: the web page footer names the model it runs on.

What we learned

  • When a requirement is decidable in Python (len(block) <= 500, "this line must be present", "no English sentence in a Korean field"), the model must not decide it — and, more importantly, must not report whether it was met. The caller re-runs the check.
  • The tool docstring is the highest-leverage prompt in a Strands agent. Field names are instructions too.
  • Deterministic checks catch what is decidable. They do not catch meaning: the Korean listing translated "BaZi" as 바지 (trousers) and passed every vocabulary rule. Any model-written copy still needs one native read before it ships; the agent's job is to make that read the only thing left.
  • Verifying that code is correct and verifying that it will ever run are separate checks (the refund handler).

What's next for ReleaseKeeper

  • Wire the App Store Connect listing caps and REFUND_REVERSED handling end to end on a second app.
  • Deploy the AgentCore contract that is already implemented, and expose the four tools as a release-pipeline step (exit 1 already makes it a CI gate).
  • Native-review queue: collect the needs_input and "native read required" items across a release into one page, so the human's part of the release is a five-minute list instead of an evening.

Built With

Share this project:

Updates

Submission history