Inspiration

Every new manager has a conversation they keep pushing to next week. Telling someone their work has slipped. Saying no to a boss who never quite hears no. Sitting down with someone whose role is going away.

The advice available is all theory: frameworks, scripts, "use I statements". None of it survives contact with a person who interrupts you, or goes quiet, or agrees warmly and changes nothing. You do not get better at those conversations by reading about them. You get better by having them, and the only place to have them is at work, where the cost of doing it badly is somebody's actual week.

So: somewhere to get the reps in first.

What it does

Pick one of eighteen situations: the original eight, plus ten for people newly managing others. Read a thirty second briefing: the setup, and two lines on what "good" looks like. Then have the conversation, out loud or by typing, with a character who behaves like the difficult person actually in front of you.

At the end you get a scorecard: four marks out of five for Clarity, Empathy, Firmness and Outcome, up to two moments you handled well quoted back from your own words, the one line that cost you with a stronger way to say it, and one thing to try next time. One, not five. A list of five things to improve is how a practice tool becomes discouraging, and nobody acts on five anyway.

The part we spent longest on is that the eight original characters are eight different kinds of difficult, not one difficult person in eight costumes:

Character The difficulty
Openly angry Talks over you before you finish the second example
Charming steamroller Never argues; absorbs the objection and starts allocating your team
Dismissive "not now" It is always the wrong moment
Evasive excuse spinner A reason for everything, a commitment to nothing
Hurt and emotional Takes it personally, and the hurt is real
Political CYA Wants it in writing, wants it to be someone else's
Cold stonewaller The curveball is silence
Blame deflector It is the process, the brief, the other team

An earlier draft of the content failed its own review for exactly this reason: every character was warm, articulate and immovable, so the app taught one skill eight times. The recast is the product.

How we built it

React Native and Expo (SDK 56), expo-router for navigation, Zustand for state. The conversation runs through a single Cloudflare Worker that holds the model key, meters usage per device with Durable Objects, and checks subscription status against RevenueCat's REST API, so the client is never trusted about its own entitlement. For paying users the model is Claude Haiku 4.5, pinned to a dated snapshot so a model change is a deliberate act rather than something that happens overnight.

For paying users, scoring uses structured outputs, so the scorecard's shape is constrained by the API rather than requested in prose and hoped for. A parse failure path still exists behind it: it shows a neutral card that says plainly it could not score, rather than a row of zeros that would read as "you did badly".

Voice is on device speech recognition in, and speech synthesis out. Recognition only starts after the device reports on device support and an installed model. If that path is unavailable, voice input is declined and typing stays fully functional, and there is no silent network fallback. The audio never leaves the phone. Only the transcribed text does, which is what the permission strings and both stores' privacy answers say.

Challenges we ran into

The model narrates itself. Nothing in the prompt asks for stage directions, but it writes them anyway. On screen they carry a lot of the character. Spoken, the synthesiser reads them aloud in the character's own voice, so the character narrates himself and the illusion dies. The fix was to parse each reply once into ordered dialogue and direction segments: the transcript renders both, the synthesiser gets only the dialogue, and the reply is stored verbatim so a moderation report carries exactly what the model produced.

Content that passed every check and could not be sent. Every scenario failed its first call with a 400: the authored character sheets ran 425 to 471 characters against a 400 character wire limit. Nothing caught it, because the content gate validates the cards in words and the API validates them in characters, and the two had never been compared. Worse, an early probe reported the contract as verified. It had been run with a hand shortened character sheet rather than a real one, so it passed precisely because it was not carrying the shipped content.

A 4 GB phone that can kill the app in the middle of a sentence. The test device is a Galaxy A05, where backgrounding the app to check notes can be enough for One UI to reclaim it. The conversation is now written to disk after every turn, before the request goes out, since the sentence is the part that took effort, and restore is offered on relaunch, never applied silently.

Accomplishments that we're proud of

  • The guarantees are tested, not asserted. "The user's line is written before the request goes out" is a property of a returned value, not a comment: the state machine emits an ordered effect list and a test asserts the order. The suite runs 589 tests across 133 suites with no test dependencies at all, on Node's built in runner.
  • Accessibility is arithmetic where it can be. A script computes WCAG 2.1 AA for 154 foreground and background pairs across both themes. Nine pairs in light and four in dark failed the first time it ran, and were fixed.

What we learned

Verifying the thing next to the thing is not verifying the thing. Both of the worst defects in this build were green somewhere: the character sheet limit passed a content gate that measured the wrong unit, and the API contract passed a probe carrying a payload nobody would ever send. Both were caught by running the real data down the real path. The first is now a check in the release gate, and the second is now a script that sends the real shipped scenarios to the live proxy.

What's next for The Hard Part

  • iOS version: submitted, and waiting on App Review
  • More scenarios, each with its own kind of difficult person

Built With

Share this project:

Updates

Submission history