One missing word can reverse a clip. One wrong file can break the handoff.
Open the current CutProof studio · Watch Joseph Ayanda's 4:37 presentation · Source and setup · Download the tested application
CutProof brings caption preparation, clip selection, source-context review and real video exports into one local-first creator workflow. Its initial target users are solo creators and small editing teams producing source-faithful interviews, tutorials and product explanations. The final editorial decision remains explicit.
Try the complete task with one click
The September 9 launch update adds Load Passage Repair example. It opens the fictional source from the presentation and selects only its first cue. Open Evidence Desk, find related passages, inspect the omitted qualification, include the full intervening context, listen, review and export the edit bundle.
The loader checks the bundled recording and captions against their expected file hashes before replacing the current source. Existing work requires explicit replacement confirmation. A canceled load, failed download, wrong file or source edit during loading does not overwrite the current project. Loading does not approve a clip, run speech inference, or publish anything.
This task uses a short synthetic recording, not a creator pilot. The public build includes the existing signed-caption comparison and Passage Repair plus the new task launcher. Implementation and packaged assets are pinned to commit 82cd57d0c005590ffbdfd80be9d5c5cf1a053e8a. The original v1.3 and earlier v1.5 deployments remain separately available.
Start with speech, not invented captions
Attach an English recording and opt into the optional speech-model download. Quantized Whisper-tiny.en runs in a browser worker to produce timestamped captions. Review and apply them explicitly, or import existing SRT, VTT or JSON captions. Audio is not uploaded to an inference service.
Evidence Desk compares an existing caption with recognition from the actual audio. In the fictional battery example, the recording says “does not last twelve hours,” while the deliberately incorrect caption omits “not.” The model flags the disagreement without receiving the reference caption as its answer. Caption changes invalidate an old speech report. The model never silently rewrites text or approves a cut. Signed comparison preserves common plus/minus signs rather than collapsing a numeric disagreement.
Find the moment without discarding its context
Local extractive ranking suggests consecutive source ranges. Boundary Lab shows nearby context, while Evidence Desk retrieves related wording from across the transcript. This is lexical retrieval, not guaranteed semantic contradiction detection.
Passage Repair addresses a specific editing gap: a matching topic sentence could be included without the correction immediately after it. The app displays small source neighborhoods, labels matched wording separately from neighboring context, and includes the whole displayed passage plus every intervening cue when expanding a cut. It does not silently stitch separated statements together.
Continuous duration limits remain enforced. Review stays pending, and stale source, selection or media actions are rejected. One neighboring cue is not guaranteed to contain all relevant context. Neither this repair nor the new loader trains the speech model or improves the separate original boundary-risk diagnostic.
Bind the handoff to the actual source
The edit ZIP and manifest contain a SHA-256 digest and byte count calculated from the attached media. Source Lock compares a supplied manifest with a newly attached file. A different export does not pass merely because its filename and duration match.
The native Python/FFmpeg renderer validates the source lock before creating outputs. Replacing a file resets editorial reviews; restoring matching bytes does not restore approval. Editing captions, selections or review states during hashing aborts the export. A canceled check produces no partial handoff.
A same-name, metadata-changed video exercises strict byte identity; re-encoding changes the digest too. A digest is not a signature, proof of authorship, evidence of consent, or certificate that captions are true. An editor can rewrite a manifest and its hash.
Deliver actual files
The browser renders captioned portrait WebM with source audio. The edit ZIP includes relative SRT/VTT captions, original source cues, manifest, readable context report, portable project, editor segments and the native H.264/AAC MP4 renderer. Reviewed-only export excludes unreviewed selections. Standalone SRT export works without attached media.
No account or inference-service key is required for the default workflow. Optional speech weights require an initial download. For local use, extract the source ZIP, serve the application directory with a local HTTP server, and open index.html in a supported browser. The source includes the disclosed fixtures and reproduction code.
Executed release and public verification
The current launch release passed 232 JavaScript tests, 13 native identity/render checks, and 74 actual local browser workflows. The browser total comprises 40 retained speech, signed-caption and Source Lock workflows, 16 retained passage workflows and 18 launch workflows. The 232 logic tests are repeated from the earlier release, not claimed as new tests. All 132 v1.3 files and 171 v1.4 files remained unchanged; the existing speech, retrieval, binding and rendering modules also remained byte-identical.
The immutable public deployment then passed 58 anonymous HTTPS browser workflows, comprising the 40 retained workflows and 18 launch workflows. These exercised actual Whisper inference, native browser hashing and media playback, source mismatch rejection, explicit edits, real downloads, and native rendering of the downloaded edit bundle. All 330 public manifest files matched their expected byte counts and SHA-256 values. No isolated browser bridge or disabled security policy substituted for these checks.
Executed native release · Public verification · Public verification run
Controlled fault cases deliberately supply a missing or corrupted example file and edit the source while loading. Happy-path application requests use actual HTTP/HTTPS. Public runs repeat release behaviors; they are not additional independent users or a user-benefit study.
Failure history remains in the repository and Actions artifacts. The launch's first run stopped at an incorrect packaging insertion point; the next encountered an existing renderer output. The harness now selects the actual outer document and uses a clean output directory while preserving historical evidence. No application safeguard or assertion was removed to obtain a pass. Earlier v1.5 verification and its failure history remain historical evidence rather than being added to these counts.
Limits and next validation
English speech support is limited to 10 minutes / 120 MB. Recognition and word timing can be wrong. Browser rendering is real-time rather than frame-exact. Desktop-editor imports and creator time savings remain unverified.
The tests use original synthetic media plus a natural-speech fixture from OpenAI Whisper's test suite. The original boundary diagnostic remains 10 of 15 risky cuts flagged, five missed, and two false alarms among nine safe cases. These small diagnostic cases are not representative accuracy; Source Lock and the task launcher do not improve that separate score.
Source Lock supports attached local files and same-origin demonstration media. For mutable URLs it identifies bytes read at check time, not a previously buffered response. Portable projects reset review on reopening but do not automatically authenticate newly attached media; use the exported manifest in Source Lock explicitly.
Next is an authorized creator trial comparing manual editing with CutProof on missed corrections, review time and usable exports. Paid setup or team workflow support is a commercial hypothesis, not existing subscriptions, customers or revenue.
Architecture and provenance
Plain JavaScript, HTML/CSS, Web Audio, Transformers.js 3.8.1, ONNX Runtime Web, a pinned Whisper model revision, Python, FFmpeg and browser automation. AI handles speech recognition; lexical retrieval finds source wording; file hashing checks byte identity; the operator makes the editorial decision.
The original application was created September 6–8, 2026; signed-caption, passage and task-launch improvements were developed September 9, with substantial AI coding assistance. This is the same CutProof project entered in AI Content Engine and AI Builders, not two separately invented applications.
The linked presentation uses Joseph Ayanda's recorded narration over prepared slides and recorded application footage with fictional examples. It is not an unedited live product session or a customer pilot. The older demonstration used stock synthetic narration. Original code is MIT licensed; third-party notices and model attributions are retained. No organizer acceptance of the presentation, field-tested time saving, general truth guarantee or prize is claimed.
Historical materials remain available: v1.3 studio, original demonstration, eight-slide v1.3 product deck, and browser deck. Those slides describe the earlier product release. The preview branch remains unmerged; portfolio production and the original deployment files were not replaced.
Built With
- css3
- ffmpeg
- github-actions
- html5
- javascript
- kokoro
- netlify
- onnx-runtime
- playwright
- python
- transformers.js
- web-audio
- web-workers
- webassembly
- whisper
Log in or sign up for Devpost to join the conversation.