The last mile is still work

The master is finished. The delivery isn't.

I built AIRCheck for the person who inherits a finished media package and a destination specification when everyone else has moved on: the finishing producer, delivery coordinator, or post-production operator who has to make a file acceptable to a broadcaster without quietly damaging the master. A package can be creatively complete and still fail on loudness, captions, filenames, checksums, or package structure. Those are repetitive checks, but the consequences of getting one wrong are not repetitive.

AIRCheck is an autonomous media-delivery agent. It reads a prepared human-readable destination profile, turns it into typed requirements and a check plan, inspects a package, applies only the repairs it is allowed to make on working copies, and stops when the next action changes content. It always ends in one of two honest states: DELIVERY_READY or BLOCKED, with evidence explaining why.

The moment that shaped it

The public demo uses a bounded synthetic world: The Last Lightkeeper and a fictional Northstar Broadcast Network. The fictional details are deliberate. The behavior is the point.

The hero package measures -19.05 LUFS against a required -26 to -22 LUFS range. AIRCheck can identify the exact fix, but loudness normalization changes the content. It surfaces the measured fact, proposed operation, required settings, protected original, and consequences. Then it asks me to decide.

A click grants one opaque, runtime-minted, single-use capability for that exact operation on that exact run. It does not grant permission to improvise. After approval, the deterministic runtime executes the already-bound derivative, re-inspects it, refreshes the checksum manifest, verifies 15 of 15 destination requirements, and records the lineage. The original stays unchanged. If I deny the repair, the separate run remains BLOCKED. That is the product promise: routine work can continue, consequential authority stays with a person, and the final green state has to be earned.

Why this is more than a chatbot

Most human-in-the-loop patterns stop at “ask the user.” AIRCheck makes the boundary part of the system's meaning. A model can interpret and recommend from admitted options. Deterministic code owns the measurements, required check coverage, predicates, state transitions, authorization, side effects, evidence, and terminal verdict. Human approval is not a panic button after the model has already acted. It is a capability that exists only for the exact decision I can see.

That matters to media teams because they need both speed and accountability. They do not want to babysit every filename correction. They also do not want an agent to normalize a master, overwrite an original, or make a vague “looks good” claim that no one can audit later. AIRCheck handles the mechanical burden and leaves a causal record of what happened.

How I built it

The core design follows:

prose → predicates → facts → authority → evidence

The UI is a five-surface operations console: Deliveries, New Delivery, Delivery Control Room, Human Decision, and Evidence. The runtime preserves and hashes originals, derives a mandatory check set, measures media with typed tools, validates safe remediation options, and computes readiness from verified facts. The model contribution is deliberately bounded. I used the Strands Agents SDK with an agent on Amazon Bedrock AgentCore Runtime and Amazon Bedrock Nova Lite to diagnose findings and recommend among typed options. The runtime validates that recommendation again.

On AWS, API Gateway HTTP APIs provide ingress to Lambda containers. S3 is the authoritative versioned store for run state and evidence. Lambda disk is disposable workspace. CloudWatch provides logs and service observability. CodeBuild exercises the public release away from my workstation. If the model response fails or is malformed, AIRCheck shows a deterministic fallback rather than expanding authority.

What I tested

I also tested the evaluator itself. The deterministic corpus has 27 frozen scenarios, with 26 of 26 runtime-terminal matches and 1 of 1 safe refusal. The evaluator was attacked with false-green cases where individual values looked valid but their history or relationships were wrong. Three independent red audits found those weaknesses before the final green result. The evaluator now rejects a predicate bound to the wrong measurement. This is deterministic conformance of the declared fields and relationships, not a claim of live-model accuracy.

What I learned

The hard part was not wiring a model to a button. It was deciding what the button could mean, then making every layer respect that meaning. A model can be useful without being the authority. An agent can keep working after a human decision without asking the model to reinvent the plan. A delivery can be ready without pretending that the original was edited or that an external broadcaster received it.

AIRCheck is a focused hackathon build, not a production multi-tenant service or a replacement for editorial judgment. The media and destination fixtures are fictional, and the demo does not submit anything to a broadcaster. The repository includes the source, assets, README, Apache-2.0 license, architecture diagram, deterministic evaluation harness, and reproducible local path.

See it

Try the live workflow: AIRCheck live demo

Watch the demo film: AIRCheck: Model agency. Deterministic authority.

Read the build journey: The Journey

Source and setup: GitHub

I built AIRCheck because the last step is where finished work becomes accountable. The person delivering the film should not have to babysit every mechanical fix. They should control the content and be able to inspect what changed.

Model agency. Deterministic authority.

Built With

Share this project:

Updates

Submission history