Inspiration
I ship a small fortune-reading app on Google Play alone. Every release repeats the same chores that are easy to describe and easy to get wrong: does the store listing still carry the AI-disclosure and "not advice" framing lines it needs? Which RevenueCat webhook event is this, and do I revoke the entitlement? Turn a 39-commit git log into release notes and a "what's new" block under Google Play's 500-character cap. Write the Korean listing I have been putting off because I cannot proofread it well.
None of these is hard. Each one costs an evening of judgment work, and skipping one gets a release rejected or a chargeback. I wanted an agent that does the work and then hands me the short list of things only I can decide — not one I have to trust blindly.
What it does
ReleaseKeeper is one Strands agent with four deterministic tools, run from a CLI or a one-page web UI on real release inputs:
| Tool | Input | Output |
|---|---|---|
disclosure_check |
store / result-screen copy | red · yellow · green findings with a concrete fix each; verdict BLOCK / SHIP_WITH_FIXES / SHIP |
refund_triage |
a RevenueCat webhook JSON body | event class, what to do to the entitlement, and a ready-to-edit reply (en / ko) |
release_notes |
git log --oneline |
New / Improved / Fixed notes plus a store "what's new" block under the Play cap |
listing_brief → listing_check |
app facts JSON + target language | a full store listing written by the model inside caps and mandatory disclosure lines, then checked deterministically |
Every tool also runs without any model (--raw), so the deterministic checks are usable on their own and testable without a key. Findings are risk flags with a fix, never legal verdicts; anything not decidable from the inputs comes back under needs_input.
The design rule that runs through all four: the generative step is bracketed. A deterministic tool sets the constraints before the model writes, a deterministic tool measures what it wrote, and the CLI / web page re-run that measurement themselves and exit 1 on a red finding. The model plans, calls and explains. It never gets to be the last word on a rule.
How we built it
- Strands Agents SDK: a single
Agent, five@toolfunctions (plain Python, docstring = the contract the model reads), a per-tool system prompt. Each request is a single turn (prompt → tool → answer) because chat-completions endpoints do not carry reasoning content across turns; for release chores that is the honest shape anyway — there is a document and a verdict, not a conversation. - Model:
LiteLLMModelpointed at an OpenAI-compatible endpoint. Default is an OpenRouter free-tier model, so the whole run costs $0; switching to Amazon Bedrock or any other endpoint is three environment variables (RK_MODEL,RK_API_BASE,RK_API_KEY), no code change. - Surface: one FastAPI process serves the web page (four tabs × "Run tool" / "Ask the agent"), the OpenAPI docs, and the Amazon Bedrock AgentCore Runtime contract (
GET /ping,POST /invocations, port 8080, ARM64Dockerfile). The AgentCore contract is implemented and verified locally; deployment was intentionally skipped for this submission (optional under the rules). - Tests: 51 unit tests, none of which need a model — because everything that decides anything is deterministic.
- Demo video: reproducible from the repo (
demo/): Playwright drives the real web page on the real inputs, ffmpeg assembles the clips; narration is synthetic text-to-speech and is disclosed as such.
Architecture diagram: docs/architecture.png in the repo (also attached to this submission).
Challenges we ran into
Everything below came from running the tools on my own live release (Saju Today 1.1.17), not from fixtures. The full list is in docs/e2e_saju_1.1.17.md.
- The model does not count. Asked for a store block under 500 characters it returned 622 with a confident "(under 500 chars)" label. The fix was not a prompt: the CLI and page now measure the last fenced block themselves, fail with exit 1, and offer a version trimmed to whole bullets. On a later run the model skipped the code fence entirely, so the extractor learned a fallback — a checker that only parses the model's tidiest output silently passes on a bad day.
- The model copies drafts. Handing it a field called
store_whats_newgot my raw commit subjects pasted into the store text. Renaming the fieldsdraft_*and saying "developer jargon — rewrite, never copy" in the docstring changed behaviour more than any system-prompt edit. - The model does not translate what it can copy. An English refund line landed verbatim inside the Korean listing twice. The brief now substitutes the language's default line, and
listing_checkreds any English sentence in a Korean description. - RevenueCat has no
REFUNDevent. A refund arrives asCANCELLATIONwithcancel_reason = CUSTOMER_SUPPORT; a handler written for aREFUNDtype never fires. The tool also warns that no refund event reaches you at all until the Play service account is registered in RevenueCat — code being right and detection being wired are two different checks. - A per-session commit history has no
feat:/fix:commits. 38 of 39 commits in the real range werechore(session); the first run produced an empty note without saying why. The classifier now has a keyword pass and names the flag to use. - Free-tier reliability. Recording the demo hit "Upstream error … Service temporarily overloaded" (HTTP 502) four times; the recorder retries with a pause. That is the real price of the zero-cost path.
Accomplishments that we're proud of
- Tested on a real release, with the numbers to show for it (one pass, default free model, wall-clock):
| Step | Real input | Result | Time |
|---|---|---|---|
check |
live Google Play listing (3836 chars, Play Developer API) | SHIP_WITH_FIXES — one real finding: no minimum-age line |
0.0 s raw · 7.8 s agent |
notes |
the real 39-commit release range | 622/500 caught by the post-check → 364/500 on the next run | ~50 s |
listing --lang ko |
app facts from the live listing (the app had no Korean listing) | full Korean listing, every field under cap, check re-run by the CLI agrees | 22–71 s |
refund |
RevenueCat CANCELLATION / CUSTOMER_SUPPORT payload (anonymized) |
class refund, revoke entitlement, reply drafted, Play-credentials warning |
3.9 s |
- A full four-tool pass in the demo recording used about 32k input / 11k output tokens in total at $0 model cost.
- 51 unit tests without a model. Public MIT repo with commits dated inside the submission period. Reused rule logic from my own pre-launch checklist is re-implemented and disclosed in the README.
- The disclosure tool discloses itself: the web page footer names the model it runs on.
What we learned
- When a requirement is decidable in Python (
len(block) <= 500, "this line must be present", "no English sentence in a Korean field"), the model must not decide it — and, more importantly, must not report whether it was met. The caller re-runs the check. - The tool docstring is the highest-leverage prompt in a Strands agent. Field names are instructions too.
- Deterministic checks catch what is decidable. They do not catch meaning: the Korean listing translated "BaZi" as 바지 (trousers) and passed every vocabulary rule. Any model-written copy still needs one native read before it ships; the agent's job is to make that read the only thing left.
- Verifying that code is correct and verifying that it will ever run are separate checks (the refund handler).
What's next for ReleaseKeeper
- Wire the App Store Connect listing caps and
REFUND_REVERSEDhandling end to end on a second app. - Deploy the AgentCore contract that is already implemented, and expose the four tools as a release-pipeline step (
exit 1already makes it a CI gate). - Native-review queue: collect the
needs_inputand "native read required" items across a release into one page, so the human's part of the release is a five-minute list instead of an evening.
Built With
- amazon-web-services
- docker
- fastapi
- ffmpeg
- litellm
- openrouter
- playwright
- pytest
- python
- strands-agents
- uvicorn
Log in or sign up for Devpost to join the conversation.