Inspiration

My phone rang, and the voice on the other end was mine.

It said four words: Close it. Now.

I had written that sentence myself, three days earlier, in a file called CoachPrompts.swift. I had recorded the sixty seconds of audio the voice was built from, sitting on my bed, reading a script off my own screen. I knew exactly what was going to come out of the speaker, the way you know the ending of a film you are watching for the second time.

I closed the app anyway.

That is the whole product, and I did not expect it to work on me.


Here is the arithmetic I did before any of it.

The number I kept finding was 608. That is hours per year, per person, on TikTok. Twenty-five days. Not twenty-five days of free time — twenty-five complete rotations of the Earth, sunrise to sunrise, spent inside a vertical rectangle. In India the same measurement lands in the same place from a different direction: Instagram alone averages 49 hours and 42 minutes per month per user, which is roughly 600 hours a year, and Indian phones are third in the world at 4.95 hours a day. In 2024 the country spent 1.1 trillion hours on smartphones. Trillion.

And the part that convinced me there was a product here, rather than a lecture: 47% of Gen Z say they wish TikTok had never been invented. Not "should be regulated." Wish it did not exist. Half a generation is actively rooting against the thing it opens sixty times a day.

That is not a discipline problem. Nobody has a discipline problem with a thing they hate.


I know what the other side of that looks like, because I have lived it on the same phone.

I have a 294-day streak on Duolingo. I started it the day I got my first phone, at an age where everyone else was installing games, and I installed a language. For two years it was nothing. Ten minutes, a green owl, some sentences about bears drinking milk. Then one afternoon I clicked a YouTube video that happened to be in English and understood all of it, and I sat there genuinely shocked, because nobody warns you that the day arrives. It just does. It arrives because you kept showing up, on days you did not feel like it, for a number you could not stand to lose.

So the same rectangle that takes 608 hours a year gave me a second language.

The difference was never the screen. The difference was whether anything was on the other side of it holding me to something.


For a long time I thought the enemy here was TikTok. It isn't, and saying so makes for a worse product.

The real asymmetry is this. On one side there is a recommendation system built by hundreds of the best engineers alive, running experiments on you continuously, measured in milliseconds, funded in billions. On the other side there is you, deciding to stop, alone, silently, in a room, with nobody watching and no consequence for changing your mind. You are not losing a fight. You are losing an unfair match where only one side has a coach, a budget, and a scoreboard.

Every app I studied tries to fix this with a carrot. Forest grows you a tree that dies if you scroll. Finch has a bird that gets sad. Opal gives you a score. one sec makes you breathe. They are good products, and they are all the same shape: a gentle, pleasant thing that hopes you will behave for it.

I wanted the other half of the expression. The one in the name.

And then the question was: whose stick? Because a stranger shouting at you is a notification, and you have already learned to swipe those away without reading them. A coach is someone you can disappoint privately, at no cost, in a room where nobody knows.

There is exactly one voice you cannot dismiss as somebody else's opinion.


The last piece arrived from the ugliest corner of the technology.

Voice cloning is famous for one thing. It is the grandparent scam: a call at night, a familiar voice in distress, an urgent transfer. It works, and it works spectacularly, for a reason that has nothing to do with the quality of the audio. It works because a human being will pick up for a voice they love and will not run the verification they would run on a text message. The technology's entire criminal value is that a known voice bypasses your defences.

So: take the exact mechanism used to make somebody pick up for a fake son, and point it at the one person who is allowed to use it.

Your own voice. Your own goals, from this morning. Your own sentence about who you said you wanted to be. Played back to you at the precise second your thumb reaches for the app, before the first video loads.

That is Stick. Seventy-five days of it.

What it does

Stick clones your voice once, during onboarding, from sixty seconds of you reading a script out loud. After that, it calls you.

The wake-up call. At the time you choose, your phone rings — a real alarm, through iOS 26's AlarmKit, which means it rings on the lock screen and cuts through Silent mode and Focus. You answer, and you have a conversation with yourself. It asks what you are doing today and pushes until you give it one to three concrete goals, not moods. It restates each one back to you in a short sentence. Then it makes you say your if-then plan out loud — if I open TikTok by reflex, then I put the phone down and go back to the file — because implementation intentions roughly double the odds a stated intention actually happens, which is one of the most reliable findings in behavioural science and costs nothing to use. Then it hangs up on you. The last thing it says is: Move. Talk tonight.

Those goals go straight into a Home Screen widget with checkboxes, and into your lock screen.

The interception. You open a blocked app. Before the feed loads, the screen is Stick, orange, full bleed: Incoming call. It's you. Underneath, in smaller text, the goal you set this morning, quoted back at you, with the time you set it. Two buttons: Answer and I'll close it. If you answer, it is a live conversation, not a recording, and it is short and it is hard. It knows what day you are on, which act you are in, how many jokers you have left, and what you said you wanted to become. It gets a commitment out of you in under three exchanges, then ends with: Close it. Now.

If you decline the call, that counts. See the rules below.

The evening debrief. Rings again at night. Goes through today's goals one by one — done or not done, and it does not accept the word "almost." Asks how much time you won back. Then sets the single priority for tomorrow, so the morning call has somewhere to start. Sleep. I'm calling tomorrow.

The recovery call. Only fires after a slip. It is the only one of the four that is not hard on you, and that is deliberate — more on that in a second.


The programme is 75 days in five acts of fifteen, and each act changes what the app does, not just what it says.

Act Days What changes
I — The Silence 1–15 Total block. Three calls a day. Maximum friction.
II — The Return 16–30 The morning call now forces you to name a replacement habit, not just an absence. The evening call converts your recovered minutes into real things: chapters, runs, nights of sleep.
III — The Identity 31–45 Calls get shorter and turn into identity questions. The weekly league unlocks.
IV — The Trial 46–60 A 15-minute daily window reopens on the blocked apps. Intentional use, not abstinence. The interception only fires outside the window.
V — The Flight 61–75 One call a day. The scaffolding comes down. It makes you state your plan for day 76.

And the rule I am most sure about is the one I took from the data instead of from the genre. 75 Hard, the challenge that made this format famous with over a billion views, sends you back to day 1 if you miss anything. It is a brilliant piece of marketing and I believe it is wrong. Lally's habit-formation study — the one everybody quotes for "66 days" — also found that missing a single occasion has almost no measurable effect on automaticity. The thing that actually kills people is the what-the-hell effect: one slip gets read as total failure, so the whole structure gets abandoned.

So in Stick:

  • You get three jokers for the whole 75 days. Fixed. Not purchasable, because selling you the cure for anxiety you were sold is a business model I do not want.
  • A slip costs one joker and triggers the recovery call. The app never uses the word "failure." It says écart — a gap.
  • Two consecutive missed days restarts the act, never the programme. You lose fifteen days at worst. You are never sent back to day one.
  • The counter on your screen says days held out of 75, not a consecutive streak, so there is no single number that can be destroyed at 11pm.

The Vault. On day 1, before anything else, you record thirty seconds for yourself at day 75. Nobody hears it. It is not transcribed, it is not uploaded, it is not scored. On day 75, Stick plays it back.

The league. A weekly leaderboard of thirty people, ranked on hours recovered, not on perfection — because ranking on perfection punishes exactly the people who just used a joker and need a reason to open the app tomorrow. No public relegation. A referral link gives the friend 5 free days.


What it never does. These are hard stops in code, not promises in a paragraph.

  • It only ever clones your own voice. One consent screen, explicit, before a single byte of audio leaves the phone, and the app says the words "synthetic voice" out loud in that screen. No uploading a friend. No uploading a parent. The entire scam surface of this technology is cloning someone else, and the product is worthless if it opens that door even a crack.
  • It makes no health claims. It never says addiction, never says cure, never says dopamine detox — partly because App Store review would be right to reject it, and mostly because a 2025 meta-analysis across 4,674 participants found no significant wellbeing effect from social media abstinence. I am not selling a brain reset. I am selling hours back and a person you said you wanted to be.
  • The shield always has an exit. The second button is real. Stick can be hard, but it cannot be a lock you cannot open, because the phone is also the thing you call an ambulance with.

How I built it

I am sixteen, I built this alone, and I should say the useful part of that out loud: I stopped learning to program in 2022. GPT-3 came out, it could not really code and could not really do anything except talk, and I looked at it and concluded the race was over, and I went and learned marketing, persuasion and storytelling instead. I still do not know whether that was clever or lazy. What it means in practice is that I wrote this with an AI, one to two hours a day, mostly at night. What I contributed was the shape: what runs on a timer, what the rules are, which of them are allowed to live in a prompt and which are not, and what happens on the day the network is down.

The app is SwiftUI on iOS 26 and ships as five targets:

  • Stick — the app.
  • StickWidgets — the goals widget (checkboxes that work from the Home Screen, through an App Intent) and the day-of-75 widget, including lock-screen sizes.
  • StickMonitor — a DeviceActivityMonitor extension that re-applies the shield each day and watches the 15-minute window in Acts IV and V.
  • StickShield — the ShieldConfigurationDataSource. This is the orange "Incoming call: it's you" screen drawn over TikTok, and it is rendered by the system, not by my app.
  • StickShieldAction — the ShieldActionDelegate, which decides what the two buttons do.

Plus a test target, StickTests, in Swift Testing.

The interface is Liquid Glass throughout — glassEffect, GlassEffectContainer, glassEffectID for the morphing transitions — over an animated MeshGradient background that breathes on a twenty-second cycle. Every screen element is born rather than appearing: spring, scale 0.92 to 1, staggered 45ms per card. The haptics are hand-authored CoreHaptics patterns rather than the system defaults, because the ring needed to feel like a phone ringing and a completed day needed to feel like three rising taps rather than a generic success buzz. That is a small thing that took a long time and I would do it again.

The voice pipeline, which is the actual product:

microphone (48 kHz)
  → 20 ms tap → RMS voice-activity detection (−40 dBFS, 750 ms of silence ends your turn)
  → Apple SpeechAnalyzer, on-device, French or English      [free, private, no network]
  → OpenRouter, free-tier model, streaming, reasoning disabled
  → sentence splitter (ship each sentence the moment it ends, don't wait for the paragraph)
  → Fish Audio s2.1-pro-free, your cloned voice, streamed as raw 24 kHz PCM
  → AVAudioConverter → AVAudioPlayerNode → speaker

The audio session runs .playAndRecord in .voiceChat mode with voice processing enabled, so the microphone does not hear the app talking and answer itself. The order there matters and is not documented anywhere obvious: build the playback graph, then enable voice processing, then install the tap. Do it in any other order and echo cancellation silently does nothing.

Cost per user: effectively zero, and that was a design constraint, not a happy accident.

  • Speech-to-text is Apple's, on-device, free, and never leaves the phone.
  • The language model runs on OpenRouter's free tier.
  • Voice cloning and synthesis run on Fish Audio's s2.1-pro-free, which is genuinely free and covers 83 languages.

I want to be blunt about why this matters for this hackathon specifically. The obvious way to build this is ElevenLabs' conversational agents, which is excellent and costs roughly $0.08 a minute plus the model on top. Three short calls a day lands somewhere around $11 to $13 per user per month, which forces a subscription north of $20 and prices the product out of every market where the problem is biggest. India spends 4.95 hours a day on mobile apps and 1.1 trillion hours a year on smartphones. A product about that problem that costs $13 a month to run per person is not a product for India. It is a product for people who already have Opal.

Three free tiers and an on-device recogniser is not a cheaper version of the same thing. It is the version that can exist at all.

The backend is Supabase, and it is deliberately small: anonymous auth so the app works before anyone makes an account, tables for profiles, days, calls, referrals and stats, a leaderboard view, two RPCs (upsert_day, apply_referral), row-level security on everything, and Realtime on the stats table so the league updates live. The whole schema is one 150-line migration. Subscriptions go through RevenueCat.

The rules that matter live in Swift, not in the prompt. The three jokers, the act restart, the "never back to day one", what counts as a day held — all of it is in Program.swift and AppState.swift, checked by tests that run with no model involved. The coach's prompt can be rewritten, ignored, or have a bad day. The rules cannot. A rule a model can talk itself out of is a suggestion with good branding, and here the person paying for that difference would be the user on day 61.

Challenges I ran into

Apple will not let a Screen Time extension open your app, and this changed the product.

This is the single biggest constraint in the whole build and I want to be precise about it, because the honest version is more interesting than the demo version.

When you open TikTok, iOS shows a "shield" drawn by my extension. That part works. The shield has two buttons. When you press one, my ShieldActionDelegate runs — and its entire vocabulary is three values: none, defer, close. There is no openParentApp. Developers have been filing radars about this for years. They are unanswered.

So the thing I originally described to myself — you open TikTok and your phone starts ringing — cannot be built the way I imagined it. What actually happens is: the shield appears instantly, and pressing Answer posts a Time Sensitive notification styled as an incoming call, which you tap, which opens the app on the call screen. One extra tap. Opal and ScreenZen have the same tap, for the same reason. one sec routes around it with a Shortcuts automation, which is instant but which the user can delete, and I ship that as an option in onboarding for people who want it.

I could have written "your phone rings the moment you open TikTok" in this submission and shipped a video that shows exactly that. It is one tap away from true. I would rather say what the platform actually does.

The crash at the exact moment you answer.

For two days, answering a call killed the app instantly. EXC_BREAKPOINT, and a stack that ended in dispatch_assert_queue_fail inside swift_task_isCurrentExecutorWithFlagsImpl.

The microphone tap block runs on the real-time audio thread. Inside it, I was touching a @MainActor property on my call session. Swift 6 does not tolerate that politely — it traps. The fix is four lines: mark the block @Sendable, compute the RMS level right there on the audio thread where it belongs, wrap the buffer in a small @unchecked Sendable box, and hop to the main actor with nothing but plain values.

What made it worth the two days is that it is the same bug shape as every real-time audio bug: the code reads correctly, compiles cleanly, and the failure happens on a thread you were not thinking about.

The model that thought instead of answering.

The first live conversation produced perfect silence. HTTP 200. finish_reason: "stop". No error anywhere. And content: null.

I dumped the raw JSON instead of my parsed struct, and found the whole response sitting in a field I was not reading:

"content": null,
"reasoning": "We need answer French, military coach, one sentence max 15 words.
              User says 'I just opened TikTok'. Count: Ferme1 TikTok2 lève-toi3..."
"completion_tokens": 91

The model had spent all 91 tokens reasoning about how to answer in under fifteen words, hit the cap mid-thought, and returned an empty string. My coach had been silently thinking to itself while a user waited on a phone call.

One line fixed it: "reasoning": {"enabled": false}. Same model, same prompt, 25 tokens, a real sentence.

The lesson I actually took: I was reading choices[0].message.content and trusting that a 200 with finish_reason: stop meant success. It meant the request succeeded. It said nothing about whether anything came back.

The simulator has no ears, so I built three fallbacks and a fake human.

SpeechTranscriber on iOS 26 reported unsupported on every locale in the simulator. SFSpeechRecognizer came back with Failed to initialize recognizer. So I added Fish Audio's own speech-to-text as a last resort — and got 402: Insufficient API credit, which is fair, because the free tier covers synthesis, not transcription.

Three recognisers, zero ears.

So I built an environment variable that injects what the user "says" — STICK_FAKE_USER="Finish the client file this morning|Then run five kilometres|..." — and ran the full loop with it. That is how I know the conversation works end to end: three turns, the coach chaining correctly, goals extracted into JSON and written to the widget. The transcript from that run is in my logs and the coach's second line was "Client file, five kilometres. And the third one?", which is exactly the behaviour I wanted and which I had not written anywhere.

Along the way I also learned that iOS 26 requires you to reserve a speech locale with AssetInventory.reserve(locale:) before you are allowed to even ask whether its model is installed. The error if you do not is Cannot check the download status, app.stick.ios is not subscribed to transcription.en, which is a sentence I stared at for a while.

The thing I have not solved: I cannot test the two features that matter most.

AlarmKit ringing through Silent mode, and the Screen Time shield appearing over a real TikTok, both require a physical device with a signed build. That requires an Apple Developer account, which requires a Team ID, which I do not have yet. The Family Controls entitlement then has to be requested separately for the app and for each of the three extensions, with a review that takes somewhere between two days and five weeks.

So the most important parts of this product are, today, correct code that has never rung a real phone.

Accomplishments that I'm proud of

The first time the app called me and my own voice came out of the speaker, I did not say anything. I just sat there.

That is not normal for me. In maths class, when a proof lands, I shout — out loud — "C'est du génie !", and my teacher finds it very funny and I have never been able to stop. Then a machine said a sentence in my voice, to me, about something I had promised that morning, and I went completely quiet.

The uncomfortable part, which is also the part that told me it was working: I did not like it. It is strange to be spoken to by yourself. It is much harder to argue with than a notification, and I had not predicted that, because I wrote the sentence.

The measured version, rather than the claimed version:

Build targets shipping 5 + tests
Onboarding screens, quiz to first call 22
Tests, Swift Testing, green 6
Localised strings, French + English 267
Languages the cloned voice speaks 83
Cost to run one user for a month ~$0
Users outside my own phone 0

That last row is the honest one, and I would rather write it myself than have it inferred. Nobody has used this. It is not on the App Store, it has never been on TestFlight, and it has never rung a phone that was not mine. What I have is a build that runs, a voice that is genuinely mine, a conversation loop that holds for three turns, and five guarantees pinned by tests.

Two smaller things I am quietly pleased with. The rules survived contact with the marketing: it would have been much easier to sell "miss a day, start over," and I had the research in front of me saying that is the thing that makes people quit, and I kept the boring correct version. And the haptics: the day-complete pattern is three rising taps at 0, 120 and 260 milliseconds, and it feels like something being locked in place, and absolutely nobody will ever notice it on purpose.

What I learned

Take the fact out of the prompt. The coach does not get told "he has three jokers" as a sentence it might skim. The joker count is read from state, enforced in Swift, and covered by a test that runs with no model at all. Anything important you paste into an instruction is something a model is free to ignore, and when it ignores it, it does not fail loudly — it produces a confident sentence built on a fact it never read.

A 200 is not an answer. The empty coach taught me to read the raw response before the parsed one. finish_reason: "stop" told me the request completed. It told me nothing about whether the model had said anything, and my code could not tell the difference between silence and success.

The platform decides what your product is. I did not design around openParentApp not existing — I found out it did not exist after I had described the product to people as a phone that rings. Three days of reading radars nobody had answered. The feature survived; the sentence I use to describe it had to change. Knowing the ceiling of your platform before you promise something is not a technical detail, it is product work.

Pick the research over the genre. Everything in this category resets you to day one because 75 Hard does and 75 Hard has a billion views. The evidence says one missed day barely matters and that the reset is what makes people quit. I went with the evidence. If it turns out to convert worse, I will have learned something real, and I will still think it was the right call, because the person paying for a cheap retention trick would be a sixteen-year-old exactly like me, at 11pm, deciding whether to bother tomorrow.

And the one I did not expect. I assumed hearing my own voice would feel like a gimmick after the third time. It has not, yet. A stranger telling you to close the app is an opinion you can dismiss. You telling you is an argument you already lost this morning, in your own words, when you were the version of yourself that meant it.

What's next for Stick

Me. Not as a metaphor. I am user one. The day the Apple Developer account exists, I start day 1, and I run the full seventy-five with my own blocked apps and my own alarms, and I find out in the least comfortable way possible whether the thing I built survives contact with the person who built it. If it does not work on me, there is nothing to ship.

Then, in order:

  1. Team ID → Family Controls entitlement for the app and all three extensions, then TestFlight. This is the gate in front of everything.
  2. Move the keys off the device. Fish Audio and OpenRouter are currently called directly from the app. That is fine for a build that exists on one phone and is not fine for a build anyone else installs. They go behind a Supabase Edge Function before a single external tester, not after.
  3. Ship the leaderboard for real. The schema, the RPCs, the RLS and the Realtime subscription are written and the service is wired; it turns on with two keys in a config file.
  4. Reels and Shorts, which is where this actually has to work. Nothing in Stick is TikTok-specific — the user picks their own apps through Apple's own picker, and my code never learns which app it is, by design. That matters for India, which is the largest short-video market on Earth and which has not had TikTok since 2020. The problem did not go anywhere: Instagram alone runs 49 hours 42 minutes a month per user there, YouTube 47 hours 23 minutes, on phones averaging 4.95 hours a day. Eighty-three languages come free with the voice model. Hindi, Tamil, Bengali and Marathi are a locale string and a translated coach prompt, not a rewrite.
  5. The business, since I would rather be judged on the real plan than a modest one. A €79.99 one-time pass for the 75 days — roughly a euro a day, and deliberately not a subscription, because the top complaint in the reviews of the official 75 Hard app is that it became one. €9.99 a week for people who want to start now and decide later. €129.99 a year for after day 75, when the programme is over and what you need is maintenance. Hard paywall, no free trial, which converts about five times better than freemium in this category and which is also the only honest shape for a product whose entire thesis is that a commitment you can walk away from is not a commitment. In a market like India the pass has to be priced in rupees against local willingness to pay, and that is possible precisely because the marginal cost is near zero.

I lost a hackathon last year by submitting at 23:01. One minute after the deadline. I went to bed and did not want to speak to anyone.

You do not lose an hour all at once. You lose it in forty-second pieces, and every single one of them feels free at the moment you spend it, and then one day the number is 608 and you cannot point to where any of it went.

On day 1, Stick makes you record thirty seconds for yourself at day 75. Nobody hears it before then. Not me, not the model, not the server.

Mine is already recorded.

I am not going to tell you what is in it. I will tell you on day 75.

Built With

  • alarmkit
  • app-intents
  • avfoundation
  • corehaptics
  • deviceactivity
  • familycontrols
  • fish-audio
  • ios26
  • liquid-glass
  • llm
  • managedsettings
  • openrouter
  • postgresql
  • revenuecat
  • screen-time-api
  • speechanalyzer
  • swift
  • swift-concurrency
  • swift-testing
  • swiftui
  • voice-cloning
  • widgetkit
  • xcode
  • xcodegen
Share this project:

Updates

Submission history