HISHIN — The Agentic Video Production Desk

Inspiration

Recording video has become easy. Finishing it has not.

A creator can record ten, twenty or thirty takes in an afternoon, but those files are not yet a video. Someone still has to watch everything, identify the strongest moments, remove failed takes and repetitions, decide how the story should flow, synchronize subtitles, apply branding, mix the audio and render the final result.

For creators, social-media teams, educators and small businesses, a surprising amount of creative work dies somewhere between recording and publishing.

HISHIN started from a simple question:

What if an AI agent could take responsibility for that production work end to end, without pretending that a language model is good at every part of video editing?

That second part became the most important architectural decision in the project.

What HISHIN does

HISHIN is a professional video-production agent built with the Strands Agents SDK and Amazon Bedrock.

A creator gives HISHIN:

  • raw video takes (and, optionally, reference photos)
  • a creative brief, and optionally the shooting script that was meant to be read on camera
  • a target duration and format
  • optional editorial constraints

The HISHIN Director inspects the available material, understands the transcripts, compares takes, identifies repetitions, selects useful segments and builds an editorial plan. When the footage itself is missing something the brief needs — a call to action no one actually said on camera, for instance — the Director can also write new narration, which Amazon Polly turns into a real voice-over, and pair it with an uploaded photo as its visual.

It then invokes deterministic media tools to perform the operations that require exact measurements: resolving timestamps, cutting footage, assembling clips, synchronizing subtitles, normalizing every clip to a consistent resolution and frame rate, processing sound and rendering the output.

A separate HISHIN Critic evaluates the result against the original brief. If the edit violates an important constraint, the Critic can return a structured revision request and the Director performs another production pass.

The final result is presented for human review rather than silently published.

The workflow is:

Raw takes → Understand → Plan → Execute → Critique → Revise → Human review → Final video

The core idea

The most important rule in HISHIN is:

AI decides what to cut. Deterministic tools decide where to cut.

A language model can reason that one take contains a stronger opening than another.

It can understand that two sentences repeat the same idea.

It can determine that a call to action belongs at the end of the story.

Those are semantic decisions.

But an exact timestamp is not an opinion.

A frame boundary is not a language problem.

Subtitle synchronization, media duration and rendering are measurements.

So HISHIN does not ask the language model to invent precise timestamps.

Instead, the Director selects semantic references such as:

take_04 / sentence_02

The deterministic timing layer resolves that selection against the actual media data:

take_04_sentence_02
        ↓
timing resolver
        ↓
00:13.482 → 00:17.921

The model owns the meaning.

The tools own the measurements.

The renderer owns the frames.

We summarize that architecture with another rule:

The agent is the director. The media pipeline is its tool belt.

How we built it

We built HISHIN in the order the architecture demanded, not the order that would have looked most impressive first.

We started with the deterministic media pipeline: ingestion, transcription (AWS Transcribe), de-rushing, tightening, assembly, subtitles, sound processing and mixing — eight independently testable modules driven directly by FFmpeg, with no model anywhere in the loop. Each one was proven against real media — real silence detection, real frame-accurate cuts, real loudness normalization — before a single line of agent code existed. That ordering was not incidental: the entire premise of HISHIN only holds if that execution layer is trustworthy on its own, so it had to earn that trust first.

Only then did we build the agentic layer on top of that engine using the Strands Agents SDK and Amazon Bedrock, wrapping every pipeline module as a typed tool — including a later addition, Amazon Polly for generated narration — that the Director can call but never bypass.

The main component is the HISHIN Director, a Strands agent responsible for planning and orchestration.

Instead of manipulating video directly, it works through structured tools that can inspect the project, retrieve transcript information, compare takes, resolve selected segments, assemble an edit and inspect the rendered output.

The second reasoning role is the HISHIN Critic.

The Critic does not edit the video. It evaluates the proposed result against the creator's brief and returns either:

PASS

or a structured revision such as:

REVISE

- Target duration: 60 seconds
- Current result: 74 seconds
- Segments 5 and 8 repeat the same argument
- The requested call to action is missing

The Director receives that observation, updates the plan and decides which tools to invoke next.

This gives HISHIN a real agent loop:

observe → reason → act → observe → revise

rather than simply placing an LLM in front of a fixed script.

Why Strands Agents

We deliberately avoided turning every stage into an agent.

Transcoding does not require an agent.

Reading a media duration does not require an agent.

Resolving timestamps does not require an agent.

Rendering does not require an agent.

Those operations remain deterministic tools.

Strands is used where reasoning and orchestration are actually valuable: understanding the creator's intent, inspecting a project, deciding which information matters, choosing tools, reacting to their results and revising the production plan.

This separation keeps the system understandable and makes agent behavior observable.

Human review

HISHIN is designed to produce a review-ready video, not to remove the creator from the process.

After production, the creator can inspect the result in a history view that keeps every version: the rendered video, the exact segments chosen with the Director's reasoning for each one, and the Critic's evaluation. Nothing is published silently, and nothing overwrites a previous attempt — every round stays inspectable.

When the Critic returns REVISE, its structured issues go straight to the Director for a new pass that must address every blocker. The number of rounds is bounded, so the system always converges on something for a human to look at, even when it still falls short of the brief.

The goal is not to automate creative ownership.

The goal is to remove the repetitive production work surrounding it.

Challenges

The rule "AI decides what to cut, deterministic tools decide where to cut" is easy to write down and much harder to hold onto once an agent can generate content, not just select it. The hardest problem was extending that same discipline to generated narration. It would have been natural to let the Director estimate how long a line of narration takes to speak — but an estimate is not a measurement, and everything downstream (subtitle timing, segment ordering, the target-duration check) depends on that number being real. So the narration tool always calls Amazon Polly first, and the agent is required to reference the id it gets back — with the actual measured duration and word-level timing already attached — rather than the text it wrote.

The second challenge was proving that boundary holds, not just asserting it. Designing a schema where an agent cannot supply a timestamp is one thing; showing that an agent which tries anyway is actually stopped is another. We built a small scripted, deterministic stand-in for the model specifically to attack the system — feeding it a segment id that was never resolved from any real transcript — and asserted that the tool call fails loudly and no file gets produced. Getting that proof running had its own detour: our test runner's module resolver couldn't follow the loader that lets our TypeScript agent code import plain JavaScript pipeline modules, so this specific test runs as a standalone script instead of inside the usual suite.

The third challenge was watching revision actually change something, on real unscripted footage. In a live run, the Critic correctly flagged that the raw takes contained no explicit call to action and that the cut fell short of the target duration. Fixing that wasn't a matter of picking different sentences — there weren't better ones to pick. The Director had to write a new closing line and generate it as narration, which is the moment the architecture stopped feeling like a demo and started feeling like it was actually deciding something.

What we learned

The biggest lesson from HISHIN is that building a useful agent is not about giving the model control over everything.

It is about giving it control over the right things.

Language models are powerful reasoning components, but reliable systems still need deterministic boundaries.

For video production, that boundary became very clear:

meaning belongs to the agent; measurement belongs to tools.

That principle extends far beyond video editing.

Many professional workflows combine subjective judgment with operations that must remain exact. Agentic systems become much more dependable when those two responsibilities are intentionally separated.

What is next

We already measure some of this — timing accuracy against the target duration, how much of the raw footage survived the cut, filler-word ratio, subtitle reading speed, and token/dollar cost per revision round — but only as scripts run after the fact, not surfaced where a creator would actually see them.

The next step is bringing that evaluation into the product itself: showing adherence to the brief, revision count and cost directly in the review view, and adding what we don't measure yet — human corrections after a PASS, for instance, which would tell us where the Critic is still too lenient.

We also want to make HISHIN increasingly useful as a persistent production desk, where a creator's brand rules and project context can survive across multiple videos without sacrificing human control.

HISHIN is not trying to replace creative judgment.

It is trying to remove everything that gets in the way of it.


Built With

Share this project:

Updates

Submission history