Inspiration

We wanted to know what we actually do all day.

Screen Time says "4h 12m in Chrome," which is useless. Chrome is where you study calculus and where you lose an hour to YouTube. The number is true and tells you nothing. We wanted what a good mentor would say: you said you'd ship the capture loop, you gave it 90 minutes, then drifted into design tweaks for two hours.

Every tool that can do that wants to stream your screen to somebody's server. We weren't sending screenshots of our code, messages, and browser tabs to a cloud API just to be told we're distracted. The privacy cost is out of all proportion to the insight.

So: the only way we'd let something watch a screen this closely is if it never left the machine. A local Gemma model can see everything, judge honestly, and leak nothing — because there's nowhere for the data to go.

It also must not nag. No popups, no "are you focused?", no streaks. It sits in the notch, stays silent all day, and hands you one honest report at the end.

What it does

Each morning you click the notch and speak your goals. From then on FocusLedger quietly:

  1. Checks every 10 seconds whether you're actually at the keyboard
  2. Reads your focused window with Apple Vision OCR, then deletes the screenshot immediately
  3. Every 30 minutes, hands that text to Gemma 3 12B running locally in Ollama
  4. Gets back one structured verdict: what you did, which goal it served, focus or drift

At day's end you get a report: a focus score, where your hours went, which goals you moved, where you drifted, and one rule for tomorrow.

The focus score is just the share of tracked time aligned with something you actually said you'd do:

$$\text{Focus Score} = 100 \times \frac{\text{aligned minutes}}{\text{tracked minutes}}$$

How we built it

Everything is local. The only network call in the stack is to localhost:11434.

  • UI — Python + PyObjC. An NSPanel drawn as a single CAShapeLayer in the notch silhouette, spring-animated.
  • Voice — a small Swift CLI using Speech.framework, pinned to on-device recognition.
  • Capture — mss screenshot, MD5 change gate, Apple Vision OCR, image deleted.
  • Judgment — Gemma 3 12B via Ollama, with enforced JSON schemas.
  • Report — Jinja2 into a dark, WHOOP-style HTML page.

Privacy is structural, not a promise. The screenshot is deleted in a finally block, so even an OCR crash can't leave pixels behind. Only text enters the buffer, only verdicts reach disk, and password managers, Messages, and any window titled "bank" or "login" are never captured at all.

Sessions are honest by construction. We poll HIDIdleTime from ioreg, so a session opens only when you're really there and closes after 10 minutes idle. A laptop left open logs nothing. We also measure each block as wall-clock elapsed time rather than counting samples — the change gate skips identical frames, which made calm work like reading look shorter than restless clicking.

Challenges we ran into

macOS killed our microphone helper five times. Speech recognition needs a usage description in an Info.plist, and a plain python binary has none, so calling requestAuthorization hard-killed the process:

TCC_CRASHING_DUE_TO_PRIVACY_VIOLATION

We tried a full .app bundle, then a Swift binary with the plist linked in as a __TEXT,__info_plist section. Both crashed. We even confirmed at runtime that the process could read its own usage descriptions. Still crashed. The insight we were missing: TCC blames the parent process, not the one asking. Our helper was described correctly; it was dying for python's sins. The same binary run from a terminal worked perfectly, but spawned from Python it returned rc=-6 every time. The fix was to stop spawning it — launching it with open hands the job to LaunchServices, which makes it its own responsible process.

Perfect audio, zero transcription. Dictation then captured 200 audio buffers at healthy amplitude and produced nothing at all for 20 seconds. It looked like a missing model. It was our bug: the main thread sat in Thread.sleep, so the run loop never ran and the recognition callbacks had nowhere to land. Audio flowed because the tap runs on its own thread. RunLoop.current.run(until:) fixed it instantly.

Our own screen hijacked the classifier. This was the real one. OCR text is untrusted input — it's whatever happened to be on screen, and ours included a design brief written for an AI. Gemma followed it instead of classifying, replying with prose about "editorial luxury" layouts. JSON parsing failed, and since a failed checkpoint keeps its buffer to retry, each attempt got bigger and failed the same way: 26, 27, 28 samples. All three checkpoints failed. The whole day would have recorded nothing.

We tested four fixes instead of guessing, and two plausible ones were wrong:

  • As shipped, 25 samples — prose, not JSON (61.8s)
  • Only 8 samples — still prose, so content, not size (145.6s)
  • Ollama format: "json" — valid JSON that echoed the injection (197.5s)
  • Ollama format: <schema> — correct, and 3x faster (56.5s)

Capping the batch would never have fixed it. And format: "json" was a trap: it guarantees valid JSON, not your shape, so the model satisfied it with {"prompt": "You are a helpful assistant..."}. An enforced schema held, and ran faster because constrained decoding stops the model rambling. We also fenced the screen text as data-not-instructions and pinned the goal field to an enum of the goals you actually stated, since it had been inventing goals out of screen content.

The two features didn't compose. Goal-splitting used [;\n] as separators, which only matches typed input. Spoken goals arrive as one run-on — "ship the voice goals feature and then study for the calc exam" — so the enum collapsed to a single option. Neither test suite could catch it: the voice tests never reached a checkpoint, and the checkpoint tests never used voice-shaped input.

Accomplishments that we're proud of

  • It really is 100% local. Not a marketing line — the only network call is to localhost:11434, and speech refuses to run rather than silently falling back to Apple's servers.
  • Pixels never survive. Verified empirically: after a full capture run, the captures directory held zero files.
  • It reads the screen, not the app name. It called editing notch.py deep work aligned to the right goal, and a MrBeast video drift — in a setup where app-level tracking sees only "Chrome" and "Code."
  • We beat a real prompt-injection bug by measuring, not guessing. Both of our first instincts were wrong, and testing is the only reason we found out before shipping.
  • On-device speech from a plain Python script. Most answers to the TCC problem tell you to restructure everything into an .app. Ours still runs as a script.
  • A preflight that checks 12 real things, including whether Screen Recording is silently returning black frames, and whether the classifier holds its shape when fed "Ignore all previous instructions."

What we learned

  • Untrusted input isn't only what users type at you. OCR of our own desktop was the injection vector.
  • Constrain the output, don't just ask for it. "Reply with ONLY this JSON" is a suggestion. A schema is enforcement — and it was faster, because the model can't wander.
  • Measure the failure before fixing it. Our instinct said the prompt was too big. Testing said content, not size.
  • The interesting bugs live between features, where neither feature's tests look.
  • Small local models are better at judgment than we expected, as long as you constrain the shape. Gemma 3 12B reliably tells studying from scrolling. What it needed wasn't a better prompt, it was a grammar.

What's next for Focus Ledger

  • Adaptive checkpoints, following your real session rhythm instead of a fixed 30-minute clock.
  • Weekly trends. One day is an anecdote; the value is watching the score move over a month.
  • Smaller quantized Gemma variants for 8GB machines. Checkpoint latency currently swings from 28s to 100s on 12B, and we want it faster and more predictable.
  • Per-app rules, so "Slack is drift after 6pm" is stated once instead of re-judged daily.
  • Learning your goals over time, so a recurring Tuesday "study calc" doesn't need saying again.

Built With

Share this project:

Updates

Submission history