Inspiration

I ship apps on my own. Three of them are live right now: an AI diary, a queue-time predictor, and a bedtime-story app for parents.

Every morning starts the same way. Open the Play Console, read the new reviews, write replies. Open RevenueCat, see what churned overnight. Open Crashlytics, hope nothing spiked. None of it is hard — that's exactly the problem. It's small, it repeats forever, and across three apps it quietly eats the hours that should go into building.

What I wanted wasn't another dashboard. A dashboard is one more thing you have to remember to open. I wanted something that does the work while I'm not looking, and taps me on the shoulder only when there's a real decision to make.

What it does

Indie App Ops Agent runs unattended on a schedule as three guards plus a weekly report:

  • Review guard (every 6h) — pulls new store reviews, classifies each one, and drafts a reply in the reviewer's own language following that app's tone guide. Routine reviews are drafted and filed silently. Only ratings of 2 or below, refund demands, or policy risks reach me — with the draft attached.
  • Revenue guard (daily) — records subscription metrics per app and stays quiet, until something breaks pattern. When trial conversion collapsed from 12% to 4%, it didn't just fire an alarm: it came with what to look at.
  • Crash guard (every 6h) — watches crash rate and DAU, and when it caught a 2.1x spike it correlated it with the release from two days earlier and named the exception. That's the part that normally costs me an hour of digging.
  • Weekly report (Mondays) — a cross-app summary that ends with "decisions worth considering," each backed by the numbers behind it.

Three principles shaped every part of it:

  1. Silence is the default. In a full run it classified five reviews — three were drafted and filed without a word, two were escalated. The work still gets done; it just stops arriving as interruptions.
  2. The model judges; the code measures. Anomaly detection is plain, unit-tested Python, not a prompt. The agent decides what a deviation means and how to say it, but it never does the arithmetic. That keeps the numbers honest and testable.
  3. Nothing irreversible. It drafts replies but never posts them — publishing requires my approval in the web console. It proposes a price change; it never applies one.

How I built it

The agent is built on the Strands Agents SDK with Claude on Amazon Bedrock, deployed to Bedrock AgentCore Runtime and triggered by EventBridge Scheduler.

One agent, five tools: get_new_reviews, get_revenue_metrics, get_crash_stats, save_reply_draft, and the one that matters most — escalate_to_human. The tools are thin, auditable wrappers around the data sources; the agent cannot reach anything else. State lives in DynamoDB (or local JSON for development), and every escalation lands both in an email and in the web console, which is the only UI the product has — deliberately, since the whole promise is that you don't sit in front of a dashboard.

It ships with three run modes so anyone can actually run it: offline (no cloud at all — rule-based fallback, so the pipeline runs in CI), demo (bundled seed data but a real Strands agent on Bedrock — this is the mode judges can run with nothing but AWS credentials), and live (real Play Developer API, RevenueCat, and Crashlytics).

Challenges I ran into

The most interesting failure wasn't technical, it was editorial. On the very first real run against Bedrock, the agent produced a beautifully written reply draft for an angry paying customer — and invented a cause for the bug. It told the user to check a setting that (a) the tool data gave it no reason to suspect and (b) doesn't even exist on Android, since the phrasing came from iOS.

It would have been easy to miss, because the draft read well. But a confident wrong answer sent to a customer is worse than "we're on it." So I made it a rule in the system prompt: a customer-facing draft may only state what the tool data actually supports. Hypotheses are valuable — they go in the escalation to me, clearly labelled, never in the text a user sees. I also moved the support contact into configuration, because when the agent had no address to point to, it made one up.

The second challenge was resisting the urge to let the model do everything. It's tempting to hand it the raw numbers and ask "is this anomalous?" — but then you can't unit-test it, and you can't explain a false alarm. Splitting measurement (Python) from judgment (the agent) made the whole thing debuggable.

Accomplishments I'm proud of

  • It's genuinely useful to me, today, against my own three apps — not a demo built for a deadline.
  • The escalation discipline holds: routine reviews are drafted and filed in silence, and only the ones carrying a decision reach the inbox.
  • 16 tests, and the entire pipeline runs offline with no cloud credentials, so anyone can clone it and see it work in one command.

What I learned

That the hard part of an agent that acts on your behalf isn't making it capable — it's making it restrained. Deciding what it must not do (post without approval, guess a cause, change a price) shaped the design far more than the feature list did. "When should this stay silent?" turned out to be the most productive question I asked all build.

What's next

Adding App Store Connect so iOS developers get the same thing, and a hosted version for developers who'd rather not run their own AWS account — the code stays open source under Apache-2.0 either way.

Built With

  • amazon-bedrock
  • amazon-bedrock-agentcore
  • amazon-eventbridge
  • amazon-ses
  • claude
  • dynamodb
  • fastapi
  • google-play-developer-api
  • python
  • strands-agents-sdk
Share this project:

Updates

Submission history