Inspiration

Most advice about difficult conversations is about whether to have them. That was never my problem. I knew what I needed to say. I had said it in the shower, in the car, in my head at 2am. Then I got in the room and it came out as "I just wanted to check in, no worries if not."

The gap is not courage. It is that the sentence changes on the way out of your mouth and you do not hear it happen. I wanted something that would let me hear it, in private, before it cost me anything real.

What it does

You pick a conversation you are dreading, or describe your own. The other person answers in character. They are defensive, or tearful, or charming, or simply unmoved, and they do not let a vague opener slide.

Three things happen that do not happen in a normal chat app.

Your hedging underlines as you type. "Just", "maybe", "I think", "sorry to bother you", "does that make sense?" are marked the moment you write them, with a running count. Nothing is blocked. You just get to watch yourself flinch in real time.

When you finally make a clear point, they say nothing. The typing indicator starts, then stops. Three real seconds of dead air. Most people fill it and talk themselves back out of the thing they just said. Sitting in the silence is the whole skill, and it is scored honestly, because the app counts on the device whether silence was offered and whether you filled it, rather than asking a model to guess.

Afterwards you are scored on four things, quoted back word for word: whether you named the specific thing, said what it cost, left them room to answer, and held the point under pushback.

Run the same conversation twice and it shows you the line that changed.

The before and after in the demo video is real and unscripted. The hedged run came back "All criteria unmet, conversation lacks clarity and firmness", zero out of eight, seven hedges. The same conversation said straight came back "User identified specific issues, impact, and deadline, and stayed firm", eight out of eight, and "They went quiet once. You waited."

How we built it

Kotlin Multiplatform with Compose Multiplatform. 4,633 of 5,198 lines, 89%, are shared between iOS and Android. The platform specific code is 220 lines of Kotlin for iOS, 276 for Android, 41 lines of Android host and 28 lines of Swift. There is no second UI.

The parts that had to be trustworthy are deterministic Kotlin in commonMain, not model calls.

The hedge detector tokenises and matches phrases against token sequences, never substrings, so "just" does not fire inside "justify" and "a bit" does not fire inside "a bitter argument". It is backed by a hand written probe corpus of ordinary sentences, half hedged and half clean, deliberately seeded with the words that break naive matching.

The silence beat is decided by the server, which holds the counterpart's state, but counted on the device. Those counts are sent into the scoring request as ground truth, overriding whatever the model thinks it saw. Silence is earned rather than random: it is only offered after you make a clear, specific point, at most twice per run, never twice in a row.

The counterpart's state is three numbers: defensiveness, how heard they feel, and whether they have conceded. The model proposes an update and the server clamps every change to one step per turn, so a hostile colleague cannot turn grateful because you asked nicely once.

The backend is a Cloudflare Worker in TypeScript in front of Groq, with its own auth, rate limiting, and a safety check on both what goes in and what comes back.

101 tests: 61 in shared Kotlin, 40 in the Worker.

Challenges we ran into

The trial the app never mentioned. The App Store sheet was offering a free week that the paywall had never disclosed. The trial length was being read from defaultOption.pricingPhases, which is a Google Play concept and is null on iOS, so it returned zero there no matter what App Store Connect said. It now reads introductoryDiscount first. Undisclosed trial terms are their own rejection under guideline 3.1.2, and we found it by watching a recording frame by frame rather than by running a test.

A run you could walk out of and lose. Leaving a rehearsal deleted it, because the session lived only in memory and nothing was written until a run was scored. Fixing that introduced a second bug: a scored run was still offered back, and finishing it twice filed two runs with the same number. Both are fixed. Both were found by watching real usage rather than by a test suite that was perfectly green throughout.

A keyboard that swallowed the exit. On iOS, Compose defaults to OnFocusBehavior.FocusableAboveKeyboard, which pans the whole surface up to clear the keyboard. The rehearsal screen keeps the keyboard open for its entire life, so the top bar went off the top of the screen, taking "Leave" and "Done" with it. Users were trapped inside a conversation with no way out. One line of configuration, found only by watching someone else use it.

Apple, three times. Rejected for a missing demo video, then for a Terms of Use link that was inside the app but not in the App Store metadata, which turns out to be a separate requirement. The app is live now.

Accomplishments that we're proud of

The silence. It would have been easy to score "left them room to answer" against a transcript and let a model decide. Instead the counterpart genuinely returns nothing, the app shows real dead air, and the thing being scored actually happened. It is the only place we deliberately made the product feel worse to use in order to make it teach something.

And the restraint. There is no share button anywhere in DryRun, on purpose. What you type is about a real colleague, sometimes by name. Transcripts and scores never leave the device, "Delete everything" means it, and the app says so on the screen where you type.

What we learned

That the feature which makes the product worth using was also the hardest to justify keeping, because it looks like nothing happening.

And that a green test suite tells you nothing about whether the app can be used. Every serious bug we shipped was found by watching a recording, not by running tests, and the tests were passing the entire time.

What's next for DryRun: Say the Hard Thing

Android is in closed testing on Google Play, running the same shared core. After that: difficulty escalation that remembers how you did last time rather than just counting runs, and hedging trends across months rather than within a single conversation.

Built With

Share this project:

Updates

Submission history