Inspiration
Creators repeat small editing decisions on every recording: trim pauses, choose between repeated takes, and crop consistently. A generic preset cannot learn the choices a particular creator actually makes. We wanted finished edits to become reusable instructions.
What it does
EditDNA learns recurring preferences from reviewed raw-to-final examples, then applies them to a new recording. It aligns corresponding speech, measures pause spacing, learns conservative first/last equivalent-take preferences, and estimates fixed centered framing. Every match and cut stays inspectable. Creators can restore takes, adjust boundaries, lock segments, undo revisions, and export a real MP4 with captions, an edit plan, and a source-hash manifest.
How we built it
A React/Vinext interface connects to a local FastAPI service, SQLite job queue, FFmpeg renderer, and local Whisper transcription. Waveform correlation finds source correspondence. Small-data statistics aggregate reviewed preferences across independent recording sessions. Visual matching compares structural edges in corresponding frames; three independent sources are required before applying a learned centered crop. No foundation model is fine-tuned.
Codex helped implement and debug the engine and editor, generate our owned demonstration corpus, write tests, investigate video timing drift and transcript-token edge cases, and prepare the release. The demo uses Deepgram Aura-2 Orion synthetic narration. EditDNA itself needs no paid API key or YouTube channel.
Challenges and what we learned
Audio alignment must be trusted before preferences can be trusted. We added ambiguity review, source-bound edit plans, immutable profile versions, revision checks, and conservative exact-transcript take matching. Video and audio cuts required different timing treatment to avoid frame-padding drift. Visual learning is deliberately bounded to fixed centered framing; it does not infer a whole aesthetic.
Results and testing
21 automated tests pass, including actual FFmpeg decode/timing, revision conflicts, undo of framing, independent-session evidence requirements, and timing-result gates. Application lint, TypeScript, production build, desktop/mobile browser checks, fresh-workspace restore, and anonymous demo access were verified.
Our five owned slide tutorials use synthesized narration and constructed editing preferences. Three source sessions train two profiles; two different sessions evaluate transfer. The regression run recovered 50/50 accepted audio correspondences within 100 ms, matched all 16 take decisions, and measured a 5 ms pause-target error versus 235–265 ms for the generic preset. Both centered-framing preferences, 1.00× and 1.10×, transferred correctly, including checks of actual rendered frames. The evaluation sources were used during integration debugging: these are small synthetic regression results, not a sealed benchmark or proof of typical creator outcomes.
A timed manual/assisted editing pilot is implemented but was not completed. We claim no measured human time savings or retention improvement.
Try it
Public demo: https://editdna-studio.adityajevoor.chatgpt.site Complete source and latest local app: https://github.com/jozai193/editdna Release downloads: https://github.com/jozai193/editdna/releases/latest
The hosted site is a recorded demonstration with real generated outputs. Select a recording/profile and click Create my edit to load the corresponding example; the applied-profile panel identifies the displayed edit. For new uploads and actual processing, clone the repository, install Python 3.11, Node.js 22.13+, FFmpeg/ffprobe, run Setup-EditDNA.ps1, then Start-EditDNA.ps1. Open localhost:3000. The first transcription downloads the speech model. Use GitHub's latest release for the newest local fixes.
Limits and next steps
Best suited to clear English speech and mostly unchanged source audio. Music, dubbing, denoising, speed changes, and inserted footage can make correspondence unreliable. Transcription mistakes can affect take decisions. Visual learning covers fixed centered framing only; check faces and edge text. Grading, graphics, narrative restructuring, and crop timing are not learned. The local edition uses one worker. Next: improve manual editing usability and evaluate with real creator footage and timed human sessions.
Log in or sign up for Devpost to join the conversation.