-
-
White: your take. Blue: the reference. Orange: your pitch corrected onto the reference's notes. Hover for the exact note and cents.
-
Drop in a vocals-only reference and your own take, either as a file or recorded straight from your mic.
-
Pick your correction style. Strong keeps more of your natural voice, and Stronger locks fully to the reference for a heavier autotune sound.
-
If the reference isn't a clean solo vocal, Unison warns you instead of silently producing a bad result.
What it is
Real autotune works fine. Unison does something different: instead of snapping your voice to a generic pitch grid, it matches your take to a specific reference vocal's actual notes and phrasing. You upload a reference vocal and your own take, and it reshapes your pitch to follow theirs while keeping your voice's actual character.
How it works
The take and reference get aligned in time using dynamic time warping (DTW), then a WORLD-based vocoder separates pitch from voice character and reshapes your pitch curve to match the reference's notes, without touching your timbre. You can pick "Strong" (natural pitch matching) or "Stronger" (locks fully to the reference, heavier autotune character).
What we learned / challenges
Our first version cut the take into individual notes, stretched each one to match timing, then pitch-shifted it. It worked technically, but sounded choppy, and big shifts gave it a chipmunk effect. We tested it against a WORLD-based vocoder approach on three different reference vocals and the vocoder won clearly every time, so we dropped the note-chopping approach entirely instead of trying to keep tuning something that was the wrong shape for the problem.
Once the vocoder version worked on synthetic test audio, testing it on real recordings surfaced two real bugs synthetic samples never would have: near the end of a take, the DTW alignment could squish the last phrase awkwardly if it didn't match the reference's timing exactly, and on quiet or breathy stretches, the pitch tracker would fail and return its lowest possible value, causing a squeaky, wrong shift. Both got traced with real frame-level data instead of guessed at, and fixed with a fallback (using the vocoder's own pitch reading when the tracker fails) and a slope-limited alignment with a fermata guard for held notes.
There's still a known limitation: when a singer's natural register is far from the reference's, the corrected voice can come out thinner, because pitch-shifting without formant correction reduces how many harmonics land under the voice's natural resonances. We chose not to build a fix for that under deadline, and documented it rather than hide it.
Built With
- fastapi
- librosa
- numpy
- pyrubberband
- python
- pyworld-(world-vocoder)
- react
- scipy
- soundfile
- typescript
- uvicorn
- vite
Log in or sign up for Devpost to join the conversation.