Inspiration
A parent posted to r/fireTV on 28 March 2026, under the title "Kid can use thier account without my permision": "It ised to need a PIN to start. now nothing. Sneak in. Turn on TV. Use at thier own will." Five months earlier another described a kids' profile set to ages 2 to 6 autoplaying, after a Cars short, footage of a frozen body. The profile was locked. The band was set. The label on the video was wrong, so none of it mattered.
In December 2025 a federal judge approved a $10,000,000 civil penalty against Disney Worldwide Services and Disney Entertainment Operations for failing to label videos it uploaded to YouTube as "Made for Kids", and ordered Disney to build a review program, "unless the YouTube Platform implements measures to determine the age, age range, or age category of all YouTube users". Two products sit in that sentence. A court had to order a studio to build the first because nobody has shipped the second. Profile Gate builds both halves, for the one box in the house nobody is standing next to.
What it does
The home screen leads with the household's own running tally, not a hero photograph: four ruled rows and their counts, and underneath them a false-positive rate that says plainly when nothing has been flagged yet rather than printing a reassuring zero. Browsing the household's titles is a plain ruled list, not an on-screen keyboard — filtering ten items with a thirty-six-key grid was theatre, so the list is what is left. When someone reaches for a title above the profile's declared age band, Profile Gate decides how hard to make the exception, from how the remote has been handled. It returns one of three states and never two: clear, unsure, or flagged. Behaviour can only restrict. The worst mistake it can make is an adult holding a button two seconds longer than expected.
How we built it
Kotlin and Jetpack Compose against Fire OS 16's API level 28. Amazon's own content://amzn_appstore/getUserAgeData provider is queried first through a real ContentResolver, and if it answers SUPERVISED with a band, that band is authoritative and nothing else is consulted. It will not answer here: its own documentation scopes it to Fire tablets, only Texas is live, and it returns FEATURE_NOT_SUPPORTED unless Amazon enables the calling app. Every other response is treated as silence and logged.
The property the whole build is organised around is that every way of not knowing lands in the same place. A thin sample window, a missed decision deadline, a model answering UNSURE, a Converse timeout and a classifier service that is not running all resolve to held, never to a pass. A classifier that silently degrades to a pass under load is the failure this product exists to avoid, and three of the 83 tests exist only to assert that an above-band title is never clear on a first attempt for any presence reading, and that a flagged decision cannot be resolved to clear under any input.
The working signal is the one sensor a Fire TV stick is guaranteed to have. PresenceClassifier reads D-pad press cadence, hold duration, dwell before commit, directional overshoot and whether search was used, from real Compose key events. Because that heuristic is unvalidated, the product publishes its own error rate rather than a confidence score: an adult can mark a refusal wrong behind a PIN, and the coverage screen then states this household's measured false-positive rate as a number, corrections over total flagged refusals. With zero flagged events it says so and states no rate at all.
The other half of the Disney order is a real Amazon Bedrock call, and it is multimodal rather than text. Each catalogue title carries one frame, sent as an image content block to Converse in us-east-1 with the title's own declared band and one question: does this frame match that label. A FLAGGED verdict escalates that title's effective band by one step, never down; UNSURE changes nothing. The model id is a chain, us.anthropic.claude-sonnet-5 then us.anthropic.claude-sonnet-4-6 then us.anthropic.claude-sonnet-4-5-20250929-v1:0, so the build upgrades itself the day access lands. The call sits behind a local HTTP service rather than inside the APK, because an APK carrying an AWS credential is a security defect.
The interaction is built for a parent glancing across a room, so the three states read as shapes before they read as colours: clear is a slow sweep of blue light across the header and no card at all, unsure is a card with a ring that fills clockwise while a finger holds OK, and flagged is the same card with nothing to press and nothing moving. Stillness is the signal. 83 JVM unit tests, every Bedrock-adjacent one stubbed.
Challenges we ran into
The vision check was dead for most of the build without ever failing loudly. The HTTP client had a 7-second read timeout, but a real Converse call through the model chain routinely takes 9 to 24 seconds end to end, and VisionSession classifies each title once, with no retry. Every one of the ten titles timed out before Bedrock's answer came back, and that answer was gone for good: the judgement stayed pinned to unsure for the rest of the session even on the calls the server had actually finished. Raising the read timeout to 35 seconds and swapping a four-slot thread pool for an unbounded one, so ten short I/O-bound calls stop queuing behind each other, is what let real verdicts start reaching the screen instead of unsure standing in for all of them.
No published study shows press cadence separates a four-year-old from an adult, so the escalate-only rule had to be structural rather than a caveat beside it. GateSessionCorrections can only set a flag on a past event, never replay or unblock it, and GateEngine never sees a raw catalogue band, only whatever ContentJudgementEngine computed, so the safety property holds identically whichever signal produced the band.
A second bug was in the honesty report itself. With the classifier service stopped, the header read "Vision: 10/10 frames checked by Bedrock". All ten calls had failed with connection refused and fallen back to unsure, and visionSummary() was counting fallbacks as answers. The gate never wavered, but the one screen meant to be the product's most honest sentence was saying the opposite of the truth. It now counts only non-unsure judgements and reports a separate "could not reach Bedrock" clause, pinned by four tests for the outage, partial-outage and not-yet-attempted cases. The model also wrapped its JSON in a markdown fence despite an explicit "reply with ONLY a JSON object" instruction, a silent parse failure that again looked like caution rather than breakage.
The identical shape of bug is still open one field over. falsePositiveRatePercent returns no rate at all when nothing has been flagged, which is the correct answer and the one the coverage screen states plainly. coveragePercent, right next to it in the same file, returns 100 when zero attempts were gate-eligible, the same vacuously-true number the vision-check fix above exists to prevent. We are disclosing this because we found it while writing this description, not because we caught it in review, and it is unfixed as of this submission.
Accomplishments that we're proud of
The classifier catches a title whose own catalogue metadata is wrong. "Sunny Meadow Friends" is declared 0 to 12 like everything else in its row, carries the darker frame, and gates anyway. That is the Disney failure mode, caught live, on a real Converse call. The coverage screen states exact counts of attempts, decisions and deadline misses, and never says "safe".
What we learned
A coverage report is product logic and needs its own network-down test, not just the safety gate underneath it. A fallback path can make a broken feature look like a careful one.
What's next
A trained model for the presence read rather than hand-tuned thresholds, a real household PIN instead of the disclosed demo value, and a run on hardware where GetUserAgeData can actually answer.
Log in or sign up for Devpost to join the conversation.