A blind viewer watching a film gets the dialogue and nothing else. Fights, faces, a city on fire, somebody quietly leaving the room — none of it reaches them. The fix has existed for decades: audio description, a voice that says what is on screen in the gaps where nobody is speaking. It is on a small fraction of titles, it costs real money to produce, and outside English it barely exists at all.

SceneSpeak generates a description track from the film itself, and speaks it on Fire TV.

It is built around one rule, chosen because it can be proved rather than claimed: a description may never be spoken over dialogue. Everything follows from that.

Tested by Amazon on real Fire TV hardware. Amazon Appstore Automated Testing installed, launched and measured the app on four Fire TV sticks — Gen 2 (Fire OS 5), 4K (Fire OS 6), 3rd Gen (Fire OS 7) and 4K Max 2nd Gen (Fire OS 8). Compatible on all four. (The first build passed only two: it asked for a newer Android than it needed. Fixed in v1.1.)

How it works.

  1. The pipeline works out when the film is talking. Subtitles are the obvious source, and on their own they are not safe — see below.
  2. The silences left over become description slots, pulled in at both ends by guard margins. The margins are deliberately asymmetric: starting a moment late just means saying less, while finishing late means talking over an actor.
  3. For each slot it samples frames across the slot with ffmpeg and asks a vision model for one line sized to that slot — the request says how many seconds and roughly how many words are available.
  4. The line is measured against a speech-timing model calibrated on the device. A line that does not fit is sent back to be rewritten shorter, not chopped: asking for a shorter sentence returns a sentence, while cutting words off the end returns "two wire-frame figures stand in a complex industrial".
  5. The Fire TV app plays the film, speaks each line at its timestamp, ducks the film to 22% while it speaks, and hard-stops the voice at the end of the slot whatever the estimate said.

The part I did not expect. Subtitles are not a safe map of when a film talks. Tears of Steel ships a music-and-effects stem alongside its full mix, so I could check. Measured against the real audio, speech runs past the end of its own subtitle cue by up to 1.07 seconds, and 13.8 seconds of the film is spoken with no subtitle at all. Planning from the subtitle file alone put 3 of 49 descriptions on top of a voice.

So the planner does not trust subtitles. Where a film ships that stem, the pipeline compares it with the full mix in the 300 Hz–3.4 kHz speech band and derives where the film actually talks. Both files contain the same music, so the music cancels out of the comparison and what is left is a voice: +10.8 dB inside detected speech, 0.0 dB outside it. (Subtracting the waveforms does not work — they are separate renders, 7 ms apart, mastered differently. The comparison has to be in the frequency domain.)

A second rule worth having. The model called a character "Vesper" in a line at 6:51. That name appears nowhere in the film's subtitles. A blind listener would have been handed a character's name that no sighted viewer gets — a spoiler delivered by the accessibility feature itself. So describe/names.py learns the cast from the film's own dialogue and sends any line that names somebody too early back to be rewritten.

Results on Tears of Steel (12:14). 78.1% of the film is describable silence. 48 slots found, 35 lines written, 578 words, 11 slots where the writer said there was nothing worth describing.

Checked on the device, in film time. The app records where the film was when every line started and stopped. Playing the whole film on the Fire OS 8 (Android 11) emulator: all 35 lines spoken, 0 still speaking when their slot closed, 0 overlapping speech detected in the audio, and the closest any description came to the next spoken word was 1.38 seconds.

Getting there took three full playbacks, and the one that went wrong is the most useful part of MEASUREMENT.md. The second run seemed to show a line taking 8.4 s in a 6.8 s gap and three lines never spoken. The first was a measurement error — wall-clock time on an emulator drifts from film time, so the app now logs film position instead. The second was real: lines start up to a quarter second late, and the player refuses to start a line that no longer fits, so the pipeline now leaves room for that.

Try it

Built With

  • amazon-appstore-automated-testing
  • android-text-to-speech
  • android-tv
  • exoplayer
  • ffmpeg
  • fire-tv
  • gemini
  • jetpack-compose
  • kotlin
  • media3
  • numpy
  • python
Share this project:

Updates

Submission history