Inspiration

We are (mostly) second-generation Chinese Americans who, growing up, could never pay attention in Chinese school. As we got older, we felt a growing sense of shame and disconnection at not being able to speak Mandarin with our families. We wanted to learn, but traditional language apps never held our attention for long, and they have little to do with the Mandarin we actually grew up around. What we did grow up around was music: 70s songs our parents played in the car, at family gatherings, on KTV. So we built Lotus Roots to reconnect with our parents and our heritage through the songs they love.

The name comes from the vegetable we grew up eating, and symbolizes our cultural roots.

What it does

Pick a classic Mandopop song from a spinning vinyl record, listen to a line, and sing or say it back. Lotus Roots will grades every word:

  • Each word gets a colour (good, okay, wrong, missing), with its Hanzi (Chinese Alphabet representation) and pinyin (latinized phonetic representation), so you see exactly where you mispronounced certain words.
  • Tap any word to practice it on its own. You hear a slowed-down spoken reference, say it back, and get scored on its sounds and its tone, with feedback on how to place your tongue, shape the ending and move your pitch.
  • Scores combine pronunciation, completeness and tone. For example:

$$\text{overall}_{\text{word}} = 0.40\,P + 0.10\,C + 0.50\,T$$

Tone carries half the weight in practice mode, because in Mandarin a wrong tone is a different word: mā (mother) and mà (to scold).

You can also select "karaoke mode", which lets you sing the desired song from start to finish with the instrumental track.

How we built it

  • Frontend: React, TypeScript and Vite. The song screen is a turntable: a three.js tonearm you drop onto the record to play a preview, with Web Audio vinyl crackle and radio static between songs. Lyrics fill character by character in time with the singer.
  • Backend: a FastAPI service that grades each recording in layers:
    1. Speech recognition with Whisper, running entirely locally: mlx-whisper on Apple Silicon, faster-whisper elsewhere. Its transcript is matched syllable by syllable against the expected lyric, comparing each initial and final.
    2. Tone checking: we track pitch with Praat (parselmouth), convert it to semitones, ( st = 12 \log_2(f / f_{\text{median}}) ), and compare each syllable's contour with the four Mandarin tone shapes, after applying tone-sandhi rules (一, 不 and third-tone sandhi).
    3. Forced alignment with Meta's MMS wav2vec2 model (torchaudio), to find where each syllable starts and ends.

We designed a song pipeline: lyrics go through pypinyin and jieba for pinyin and word boundaries, the tracks are aligned to the lyrics, and reference clips for every word are generated with neural text-to-speech. With our group of four people, we carefully planned and wrote binding interface contracts up front (shared data types, function signatures, scoring rules), so each person could build and test their layer alone. Everything runs on laptops either Docker or a native setup.

Challenges we ran into

  • Whisper was too helpful. For full lines, mistakes on syllables are autocorrected. This goes against the goal of detecting and analyzing mispronunciations. We ended up implementing a separate analysis when you click on a syllable, which utilizes a separate model on top.
  • Whisper hallucinated English. For example, the Chinese character 好 (hǎo) came back as "How". We fixed it by forbidding every token containing a Latin letter.
  • Whisper was slow. At times, whisper would take up to 5 seconds to analyze a 1.4 second audio clip. We ended up utilizing a special model built specifically for Apple Silicon chips, which sped up analysis of that same clip to 0.6 seconds.
  • Tones are hard to hear by machine. Our tone checker worked well on synthetic voices but was close to chance on our own spoken lines, and a completely robotic, flat reading still got full marks. We chose to grade tone only where it is reliable, single words in Word practice, rather than show learners confident wrong marks.
  • Singing erases tones. When you sing, the melody replaces each word's tone, so a karaoke app can't grade tone on sung lines at all. That is why practice and feedback happens in the spoken, practice mode.
  • Honest feedback. Early versions of our project told you what you "said" ("sounded closer to l"), and that guess was sometimes wrong. We switched to hints on how to say the target sound, which helped even when the grader misheard.

What we learned

So much about Mandarin phonetics! Initial and final sound, why zh and z are different sounds, how tones used twice in a row have different sounds, and why a third tone inside a word only falls. We also learned how speech models actually behave, and that knowing what not to score matters as much as scoring.

Accomplishments that we're proud of

We're super proud that we were able to scrape together a complex system in the allotted hacking time! Our system had some complex pipelines that required careful planning, and we were able to effectively divide our work among 4 team members to achieve all of our parts of our MVP + some stretch goals!

What's next

In the future, we want to record labelled takes from real learners and native speakers, to tune the tone checker on human voices. In addition, we want to create a singing-accuracy mode with rhythm scoring (currently, we have an accuracy scoring only when learning the songs, which places more of an emphasis on correct tones vs. intonation). We also want to expand our catalog of songs.

Built With

Share this project:

Updates

Submission history