Inspiration
I build indie apps — HealthBuddy, a workout tracker with AI rep counting, is the one I've spent longest on — and for a long time I filled the 100-character keyword field the way most indie devs do: with terms that sounded right. When downloads stayed flat I couldn't tell which of three very different problems I had: the wrong keywords, being invisible for the right ones, or an app people didn't want.
I used a keyword tool. It showed me rankings, and it could tell me a rank had moved. It could not answer the only question I actually had: did that make me any money? Nothing in the ASO tool knew my revenue, and nothing in my RevenueCat dashboard knew my rankings. The two halves of the answer sat in two browser tabs and never spoke to each other.
That's the gap Shipy exists to close. Not "another keyword tool with a nicer chart" — a tool that joins ranking data to revenue data, and then writes what it learned back to the App Store.
The second push: ASO work is bursty. You grind on metadata for a launch or a big update, then you don't touch it for two months. Most ASO tools charge every month for a thing you use in bursts — a pricing model designed for agencies and inherited by everyone else.
What it does
Shipy is App Store keyword research for indie developers, native on iPhone, iPad and Mac.
For any search term it answers three questions: how many people search it, how hard it is to rank for, and where your app sits right now. It tracks that daily so you can see whether the metadata change you shipped on Tuesday actually moved anything.
Then it goes past where other tools stop:
- Revenue-attributed ASO. Connect your own RevenueCat account with a read-only key and Shipy overlays your actual daily revenue on your keyword rank history, on one shared time axis. Your rank for "habit tracker" climbed 14 places in March — here's what your revenue did that month.
- It writes back. Connect your own App Store Connect key and Shipy reads your live listing, grades it, and pushes an approved keyword field back to Apple. The research doesn't die in a spreadsheet.
- Metadata health score. Apple indexes your name, subtitle and keyword field as one bag of words, so a word you repeat across them silently burns characters you paid for. Shipy finds every duplicate, prices it in wasted characters, and drafts the rewrite.
- Competitor x-ray. Reverse-engineer the terms a rival ranks for, and find the gaps.
- Your real funnel. Impressions, product page views and downloads from your own App Store Connect analytics, split by search, browse and referrers — so you can tell a discovery problem from a conversion problem.
- It tells you what to do next. One ranked plan across every app you track, rank-alert push notifications, and a "What changed" inbox that says when each number was measured. Plus a home-screen widget and Siri shortcuts.
- It talks to Claude. Shipy ships an embedded MCP server on macOS. An agent in Claude Code, Claude Desktop, Codex or Cursor can research keywords, grade your live listing, draft a packed keyword field and read what your keywords earned — Shipy stays the source of truth, and publishing stays with you.
Pricing follows how the work actually happens: 10 keywords free for life, one-off credit packs ($1.99 for 10, $3.99 for 25 — they never expire) for a launch week, and Shipy Pro ($4.99/month or $49.99/year) only if you want unlimited.
How I built it
Two Swift codebases and no JavaScript anywhere.
The client is SwiftUI for iOS and macOS, structured as a 14-module SPM package with the Xcode
target as a thin shell. Modules are layered strictly — pure domain models at the bottom, then the
SwiftData cache, networking, design system, routing, then features on top — and a dependency is
never allowed to point backwards. Every screen is a small View plus an @Observable ViewModel
nested as an extension of that view, so the view holds layout and nothing else. Dependencies resolve
through FactoryKit, which means every ViewModel is testable by registering a mock API on the
container and asserting state transitions. Async content is modelled as an explicit
Loadable<T> — idle, loading, loaded, failed — rendered through one shared view, so loading and
empty and error states look the same everywhere and an error can never quietly become an empty array.
The server is Vapor and PostgreSQL, layered controller → use case → service → repository, with controllers kept to about ten lines. The keyword resolve pipeline is the core of it: quota gate, normalise and dedupe, load the shared cache in one query, decide what actually needs refreshing, then fan out to the upstream data providers under a bounded concurrency gate, compute difficulty and position, persist, append a history point, return.
The keyword cache is shared across all users, keyed by term, storefront and platform. One person researching "gym log" warms it for everyone, which is what makes a free tier affordable.
Difficulty is a model of how entrenched the apps currently ranking for a term are — a term defended by established apps with deep track records scores high, one where the incumbents are shallow scores low. It also reports a confidence value, so a term with three results doesn't pretend to the same certainty as one with two hundred. I tuned it against an established commercial tool across a sample of terms and landed at a mean absolute error of about 2.3 points on a 0–100 scale, which was good enough to trust and honest enough to publish.
RevenueCat does the monetization, and it does more work here than a paywall usually does. The SDK owns entitlement on the device. A server-to-server webhook keeps a projection of entitlement in Postgres, which is what the server-side gates read. And a REST fallback covers the gap between "purchase completed on the phone" and "our server heard about it", because the one thing I refused to ship was a paying user getting blocked by my own webhook lag.
Testing is about 1,180 unit tests on the client (Swift Testing: business logic, ViewModel transitions, auth token single-flight refresh, transport retry), a UI-test suite driven purely by accessibility identifiers, and a database-backed integration suite on the server for the paths where correctness is money: quota races, webhook ordering, entitlement transfer.
Challenges I ran into
A popularity number that lies quietly. The hardest problem in Shipy wasn't fetching search volume — it was discovering that one of my inputs returns the same bottom-of-scale value for two completely different situations: a term almost nobody searches, and a term it simply has no answer for. The response looks identical either way. About a quarter of a broad sample was affected. One term I knew perfectly well was heavily searched came back pinned to the floor.
That's the worst class of data bug, because nothing errors. You get a plausible number, you rank your keyword list by it, and you confidently pick the wrong terms — in a tool whose entire job is telling you where to spend one of your hundred characters.
The fix was to stop treating popularity as a single lookup and start treating it as a resolution problem with a trust hierarchy: several independent sources, ranked by how much each deserves to be believed, with strict rules about which one is allowed to override which. Every stored value carries a tag recording where it came from, so a lower-confidence source can never silently launder itself into looking authoritative. And if no source can price a term honestly, the field stays empty and is marked stale. It never gets a made-up number.
An upstream that caches its own rejections. One of the APIs Shipy depends on answers a throttled request with an empty error — and its CDN then caches that error against the exact request signature. So the textbook fix, retry the same request with exponential backoff, fails for minutes on end while a trivially different request succeeds instantly. It took an embarrassingly long time to see, because every retry was correct and every retry failed. Retries now vary the request enough to miss the poisoned cache entry, with the response fed through a stale-cache fallback so a user sees last-known data marked stale rather than an empty screen.
The RevenueCat event that has no user id. When an anonymous user signs in, RevenueCat fires a
TRANSFER event to move their purchase onto the account. That event carries transferred_from and
transferred_to — and no app_user_id. My DTO required one, so decoding threw a 400, and a webhook
that returns 400 is marked permanently failed and never redelivered. That is the normal purchase
path in a sign-in-gated funnel, so I had quietly built a system where every real purchase lost its
entitlement. It's now handled explicitly: grant to every target id, tombstone every source id, and
order everything by event timestamp so a redelivered old event can't resurrect a revoked
entitlement. There's an integration test that replays the raw sandbox payload, because I never want
to rediscover this one.
Giving away 260 free keywords a year without noticing. The free tier was originally a weekly pool. But a keyword, once researched, keeps refreshing daily for free forever — that's the retention mechanic, you come back to watch your ranks move. A renewing weekly pool meant an unbounded set of keywords being refreshed daily against a rate-limited upstream, forever, for free. I moved it to a lifetime pool, which caps the cost permanently and left the retention loop completely intact.
App Review asked me to prove it works. My first submission came back not rejected exactly, but held for information: they wanted a screen recording on physical hardware, the device list, the external services, the regional differences. Answering it properly meant auditing my own app as a stranger would, and I ended up cutting an operator-only debug screen out of the shipping build because it had no business being in a consumer binary, even unreachable. Then, while recording the demo on an iPhone 12 Pro, I found a layout bug I'd never seen on the simulators I develop on: the metric tiles were narrow enough that the value wrapped, so a difficulty of 12 rendered as "1" over "2". Not a cosmetic bug — a wrong number on screen, in a tool people are supposed to make decisions with.
My ASO tool's own ASO broke. In September I rebuilt Shipy's own listing around the highest-popularity words I could find. Impressions went up about six-fold — and the share of searchers who tapped fell from 3.3% to 1.0%, with zero downloads. One word I'd added, "position", turns out to return sex-position apps in the App Store; others were people looking for Apple's own TestFlight and Xcode. Dropping "ASO" from the name took my rank for "aso" from #16 to #232. It was the most useful thing that happened to the product: popularity and difficulty are only half the model; the other half is whether the search results page actually belongs to your category. Shipy's keyword suggestions now keep to terms that are on-topic for your app. Version 1.1.0 puts the listing back on the words people tap, and I published the numbers.
Accomplishments I'm proud of
The honesty rules, mostly. When a data source is down, Shipy says so and marks the affected values stale rather than showing you yesterday's number as if it were fresh. When it can't price a keyword, the field is empty. It never invents a value to fill a chart. That sounds like a small thing until you notice how many tools in this category won't tell you which of their numbers are real.
The revenue overlay is labelled as correlation, not attribution, and I'm proud of that too. Apple exposes no per-keyword download source outside Search Ads, so nobody can honestly claim causation here. I show the two series on a shared axis and let you draw the conclusion. It would have been easy, and dishonest, to call it attribution.
Also: it's genuinely native. One SwiftUI codebase, a full sortable keyword table on Mac and iPad, a fast one-handed list on iPhone, no web views anywhere.
What I learned
That a data pipeline's failure modes matter more than its happy path. The interesting engineering in Shipy isn't fetching a number — it's the layered fallback, the stale marking, the self-healing retry TTL, and the decision to show nothing rather than something wrong. I now think of "what does this do when the upstream is down" as a design question, not an error-handling detail.
That billing is a distributed systems problem wearing a friendly SDK. The webhook is asynchronous, the client knows things the server doesn't yet, events arrive out of order and get redelivered, and identity changes underneath you when someone signs in. Every one of those is fine on its own. Getting them all right at once is where the real work was — and RevenueCat's event model is what made it tractable rather than terrifying.
And that dogfooding your own analytics tool is a strange loop. I use Shipy to pick Shipy's keywords — and it was Shipy's own analytics that showed me my listing was attracting the wrong searchers.
What's next
Push rank alerts, a home-screen widget and Siri shortcuts shipped during the event. Next: buying Pro without creating an account first, honest per-storefront labelling wherever a popularity value is a market-wide estimate rather than that country's own number, and closing the remaining accuracy gap on hard-to-price terms.
Built With
- app-store-connect-api
- aso
- docker
- fastlane
- fluent
- icloud-keychain
- indie-dev
- ios
- jwt
- keychain
- macos
- model-context-protocol
- postgresql
- revenuecat
- sign-in-with-apple
- storekit
- swift
- swift-charts
- swift-concurrency
- swift-testing
- swiftdata
- swiftui
- vapor
- xcuitest
Log in or sign up for Devpost to join the conversation.