-
-
Seven seconds of a film with no description. Something happened. If you cannot see the screen, that is all you get.
-
The app on a Fire TV, opening a clip nothing had ever described. It offers to write one, and says so aloud.
-
Finding the gaps first: 15 places where nobody is speaking. It only writes into silence it already has.
-
The same passage, now described. Written from what changed on screen, placed where no one is talking.
-
Descriptions are measured to fit the gap before they are spoken, so they never run over the dialogue.
-
One press moves description to a phone. The room hears the film untouched; one person hears the description.
-
The phone following the television over the local network, showing every description as it plays.
-
Asking anything, mid-scene. The answer draws on the frame on screen and on dialogue from minutes earlier.
-
Measured across three kinds of video, including one where it correctly produces nothing at all.
-
A figure a few pixels across, and it is the shot. Nothing about it is in the dialogue or the soundtrack.
-
Everything is open: a live demo needing no Fire TV, the full source, and a separate MIT QA harness.
Inspiration
Audio description is not a new idea. Someone watches a film, writes down what happens on screen, and a voice reads it out in the gaps between dialogue. It works, blind and low-vision viewers rely on it, and there is a whole craft around doing it well.
The catch is that a person has to write it and someone has to pay them. So it exists for a slice of what gets made, and for everything else there is nothing: you hear a door, a change in the music, a small sound, and you know something happened but not what.
I did not want to guess at what would help, so I asked. Reviewers on the ACB Audio Description Project mailing list answered in detail, and almost none of the rules this thing follows are mine. One of them is the reason it says nothing at all during dialogue. Another is the reason it never says "the camera pans."
What it does
Sightline writes the audio description that nobody was going to write, and speaks it on a Fire TV.
It finds the gaps where no one is speaking, looks at what changed on screen across each gap, writes only the part you would be lost without, measures the line to make sure it fits the silence it has, and speaks it there.
Three things follow from building it on a television rather than as a phone app pointed at a screen:
- It fits the film. Every line is timed against a real gap. Nothing runs over dialogue, because the gap is where the writing starts.
- Two people can watch one screen. One press sends the description to a phone. The room hears the film untouched; one person hears the description.
- You can ask it things. Mid-scene, by typing or by speaking with the keyboard's dictation key: "why does she care about this creature?" It answers from the frame you are on and from the dialogue minutes earlier. A pre-recorded description track structurally cannot do that, because it was written before you had the question.
How we built it
The pipeline, per video. Amazon Transcribe gives word-level timings, which become the gaps. Frame pairs from each gap go to a model, which writes a candidate line. Lines are ranked, because at 1.5x speed there is less room and something has to be dropped first. Amazon Polly (generative) speaks them, and the measured duration is checked against the gap before the line is kept.
The app. React Native for Vega (react-native-kepler ~4.0.0, RN 0.83, SDK 0.24.9914, CLI 1.3.4). Video through W3C Media on the device, description audio through a separate PCM playback stream so it can duck the film rather than fight it.
The second screen. A small local service the phone follows over the network. The television posts its playhead; the phone plays the matching description. Asking runs through API Gateway and Lambda so it works for a judge with no local setup.
Audio analysis, verified rather than tuned. The onset classifier's threshold is not a number I liked the look of: it is derived from the measured noise floor. A blind reviewer proposed a test I had not thought of. Attenuate the whole file, and check the classification does not move. It exposed a real bug. The loader was reading 16-bit, this film decodes to samples above full scale, and those were being clipped. Reading 32-bit float fixed it, and gain invariance went to exactly 0.00 dB.
Challenges we ran into
Amazon Bedrock was refused at the account level. The client is written and shipped. Every region, every model, bare IDs and both inference-profile forms: Error 002: Access to Bedrock models is not allowed for this account, while S3, Polly and Transcribe worked on the same credentials. get-foundation-model-availability reports every field positive: authorizationStatus AUTHORIZED, entitlementAvailability AVAILABLE, regionAvailability AVAILABLE and agreementAvailability AVAILABLE. The account is not in an AWS Organization, so no service control policy explains it. The control plane says access is granted and the data plane refuses, which leaves a developer nothing to act on. Re-checked 23 September 2026 across five regions.
The platform cost us time in ways worth writing down. Fifteen findings are in FRICTION-LOG.md, written while building rather than remembered afterwards. One of them: TVEventHandler from react-native exists on the platform, registers nothing, and fails silently, so no button did anything at all.
The bugs a blind tester found, and I did not. A blind reviewer watched a three-minute clip and told me description stopped after two minutes. A limit written to cap the number of descriptions was truncating the film. On a feature-length film it would have described the first minute and then gone quiet. He also found a scene where the software reported someone watching a character sleep. Nobody was there: it had turned a camera angle into a person.
And one nobody would have reported, because it looked fine. The companion service implemented GET, POST and OPTIONS, so it answered 501 to every HEAD. The phone probes each clip with a HEAD request before playing it, read "not ok" as "neither file exists", and played nothing. Silently, with the position counter ticking over perfectly. Static hosting answers HEAD for free, so the public pages were always fine and only the co-viewing path, the entire point of the phone, was mute.
Accomplishments that we're proud of
It was tested by the people it is for, and it changed because of them. Two blind reviewers, quoted directly in the repo, including the six faults one of them found.
It knows where it stops working, and says so. Measured across three kinds of video: an animated film at 4.5% dialogue produces 24 descriptions in a three minute stretch; a 1951 instructional film at 77% narration produces 11, in the gaps that exist; a one-minute advert produces none at all, because it fills every second it paid for. That last row is in the demo video. A tool that claims to describe everything is lying about at least one of those.
The claims are reproducible. Two of the numbers in the README come with the commands that regenerate them.
It works without a Fire TV. The live demo runs the same timeline and the same generated audio in a browser, so a reviewer can hear it in about ten seconds.
What we learned
Silence is a feature, and it is the hardest one to get right. The first version described too much. The reviewers' notes all pointed the same way: nothing during dialogue, nothing that repeats what the sound already tells you, nothing about the camera. What is left is smaller and much better.
"It looks fine" is not evidence. Three of the worst bugs in this project presented as working software: a truncated film that played normally, a phone that tracked the playhead perfectly while speaking nothing, and an app that went completely silent in the one state where a blind user has nothing else to go on. The offer to describe an undescribed video was never audible, because that code path returned before the audio stream was initialised.
Ask the people who will use it, early. Every design decision that survived contact with reality came from a reviewer, and every one I invented on my own got corrected.
What's next for Sightline
- Let people bring their own video. The pipeline is per-video already; the missing piece is a safe public ingest path rather than new capability.
- Voice input on iOS. iOS Safari has no
SpeechRecognitionat any origin, so in-page voice there means recording audio and transcribing it. Transcribe runs as a batch job through S3, which is far slower than anyone waits for an answer about the shot they are on. Streaming transcription is the fix. - More blind testers, on more kinds of content. Two is not enough, and
STATUS.mdis honest about what has not been tested yet. - Keep the QA harness useful to other people. The checks that catch invented characters and camera language are released separately under MIT as audio-description-qa, because anyone writing description automatically will hit the same faults.
Built With
- accessibility
- amazon-api-gateway
- amazon-devices
- amazon-polly
- amazon-transcribe
- amazon-web-services
- aws-lambda
- boto3
- css
- ffmpeg
- fire-tv
- github
- html
- javascript
- kepler
- media-source-extensions
- numpy
- python
- react-native
- typescript
- vega-os
- web-audio-api
Log in or sign up for Devpost to join the conversation.