Inspiration

Clipping is the tax on long-form. Six decent clips from a 45-minute episode, clipping them, writing hooks, titles on three platforms, timing and posting the whole thing takes four to six hours. Each week. Always. It's the process everyone skips or rushes and which determines everything getting seen at all.

The logical solution would be to ask a language model to "find the good parts." But the language model asked to find the good parts will return the sections which summarize well, and a good summary is precisely the element removing the need to see it at all. Hooks and highlights are two entirely different things.

And that's why I wanted to try the opposite approach: Stop looking for what is good in the clip and start measuring it instead.

What it does

You provide HOOKLINE with a transcript, SRT, WebVTT, a copy-pasted YouTube transcript, or plain text, and it gives you a ready-to-publish week:

Clips ranked by their potential and precisely cut where topics change in reality, non-overlapping. New hooks based on what the speaker has said. Titles tailored for each platform with a fit inside their character limits. Captions time-coded to the word and optimized for a vertical safe area in SRT and VTT format. Hashtags generated from phrases used by the speaker. Thumbnail ideas as proof-of-concept designs with overlay text and the exact frame to use. Chapter markers for the source video, using the same boundaries as the clips. Posting schedule with the reasons each post was selected. Also, an ffmpeg script to produce the videos with captions burned in.

For a sample of 10 minutes, it finds 1,391 candidate windows, cuts 6 clips, produces 24 titles, creates 72 caption cues, and 6 chapters plus generates 31 files in 76 milliseconds, all in the browser tab.

It operates without an account, any uploads, or API key required. Your transcript stays on your device all the time.

How we built it

The nucleus is an attention curve derived directly from the transcript, by way of three stages.

First, eight signals are quantified for every sentence: Hook (20% of the score), Curiosity (16%), Salience (15%), Emotional Intensity (13%), Concreteness (12%), Payoff (9%), Quotability (9%), and Delivery Speed (6%). All are independent, identifiable measures, not an embedded vector. Discourse management is penalized rather than rewarded, since "let me be specific about what I mean" is short enough and clear enough to dupe a naïve scoring system, but is stage direction, not content. No delivery speed measure is awarded when no time codes are available, and no guess at a timing provided.

Second, topic boundaries result from lexical cohesion. A window of sentences moves along the transcript, followed by another one. For every gap between them, I calculate how similar their vocabulary is. When the similarity plummets, the subject has shifted, which is the point where the boundary must go. Only boundaries mark the start and end of a clip, which is why a clip cannot end in the middle of a thought. The same boundaries determine the chapters, so there is no way for the two results to contradict themselves on topic beginnings.

Three, all valid windows are then scored and condensed so there is no overlap at all, weighting one-third on the opening, one-quarter on the body, and the rest equally divided amongst payoff, containment, and topic boundary fitness. There is a deliberate weighting towards the opening because, on a vertical scroll, the first line must work or nothing else will be visible. Containment scores the clip negatively if it begins on a pronoun with no referent; it is a better killer than a bad hook.

The engine is roughly 3,200 lines of TypeScript code, with no runtime dependencies, no DOM manipulation, no filesystem accesses – and so identical across both the browser workspace and the Node CLI interface. And that is also why it is reproducible: the transcript generates the same clips identically on all machines.

The interface visualizes the measurement, not the other way around. The 3D landscape's heights and temperature are direct outputs from the curve analysis, as is everything it shows, which is additionally outputted in text and a 2D chart form so that it stands up even when all motion is stopped.

An optional Claude pass is capable of rewriting hooks and titles by simply providing a key. However, this is completely additive because the deterministic outcome has been determined prior to involving any model since the clips' selection and timings and captions have been measured and will never go through the rewriting process. A failure in the pass will cause the use of the deterministic text.

Challenges we ran into

Nearly everything I generated originally was always saying something no one ever said, and every instance of such a generation checked out type-wise:

The first would add a question mark to the end of an incomplete statement and generate "If you take the total cost of that, including the?" The second would join the top two highest-ranking words to form a phrase and generate "Better Hundred:" as a title prefix, and "Clip Seconds Video: What Actually Works" as a channel title, neither of which anyone ever said. The third would recycle a common phrase "here's what actually happens" and use it in three different clips of the same set, promising something that the clip might not even deliver.

All the corrections were incorporated as rules rather than band-aids: the sentence compressor must flag the fact that the sentence is incomplete to prevent adding punctuation that would make it look complete. A multi-word phrase must be confirmed as a real statement someone made, never compiled by ranking words individually. And canned statements must be eliminated because the actual statement is the bait.

There were always some problems extracting chapter titles from common conversations sentences since none of four trials ended up with anything but grammatical trash, like "If It Is Already Known". In the end, I abandoned the idea of extracting sentences and constructed labels from phrases used by the speaker, basing them on anything that could distinguish the phrase from the rest of the recorded conversation.

Then there was this WebGL hero that was silently rendering nothing. The size of the canvas was stuck at the default browser's minimal size since 3D library waits for browser's resize event which never triggers, while camera only oriented itself inside the render loop that did not run without animation frames. There is neither any error nor exception – just a black rectangle. Both are fixed from the first frame.

In addition, CI found some flaky tests which I wrote myself. I put a performance guard with 8-second budget which ran successfully in 3.5 seconds on my own computer, but failed after 8.2 seconds on a shared runner, which actually measured the level of its busyness. This became a ratio checked against a valid transcript of the same size on the same machine.

Accomplishments that we're proud of

It can't fail cold. No key, no network, no setup. Click on the link, and you'll get a complete analysis in under a second.

It's inspectable. Every snippet includes the eight signals which led to its selection, and the five parts of its score. There is no such thing as a black box here.

It's fast where it really matters. It analyzes two-hour transcripts with over seventeen hundred sentences in under 300 milliseconds.

And there are 59 tests that gate every deployment, with one invariant this entire product stands or falls by: the generated hook should not contain even a single word not spoken by the speaker.

What we learned

Type safety means absolutely nothing when it comes to generated text being truthful. All the forgeries above compiled perfectly fine and none would have been stopped by the compiler. None of them got caught until they were run on an actual transcript and their actual output was examined, no exceptions.

"Don't fabricate" needs to be an enforceable rule rather than a mere intention. This is why it became an actual test to verify all the hooks generated against what actually exists in that particular clip, rather than a comment saying not to do it.

The process of writing this test suite even resulted in finding an actual bug - a very small YouTube transcript of two or three lines in length was treated as plain text and its actual timestamps were silently replaced with an approximation based on word-per-minute calculation.

What's next for Hookline

Clip editing, so you can tweak the start and stop times of a clip, or toss out an entire clip and choose a better alternate. It's guaranteed that there will always be one clip that the creator disagrees with, and there is no current way for them to make this known.

Speaker diarization, for clips of interviews and panels so that they can be cut up according to each individual speaking voice.

Direct editing of video within the browser itself, to move from timecode input all the way to final vertical video without leaving your tab.

And eventually, a feedback loop based on actual retention data. Right now, the weighting of each signal is done by hand and I would back each one, but it ought to be done automatically by which clips retained users' attention.

Whereas the project ends at the moment, it analyzes text. It does not actually cut videos but provides timecodes, subtitle files, and metadata for further use by an editor or the mentioned ffmpeg script. Moreover, it cannot turn poor recordings into good ones. In case someone recorded an hour of silence, no amount of cutting will produce something which wasn’t actually said because it only amplifies the data and does not create it. Posting recommendations provided by the software also fall into the domain of platform conventions and represent a reasonable default option to be changed. Also, transcripts used as input data for the software are shipped with it.

Built With

Share this project:

Updates

Submission history