Inspiration
Every musician knows the gap between practicing and playing. Scales and etudes are very repetitive, boring to practice, and give you almost no signal about whether you're improving. Additionally, popular rhythm games will reward you for tapping plastic buttons in time, not for making sound. We wanted the loop that makes roguelikes addictive (climb, fail, learn, climb again) directed to something that actually makes you better: putting the horn to your face and reading something you haven't seen before.
Slay the Spire gave us the shape. Sight-reading gave us the cards.
What it does
Slay the Choir is a roguelike where every card is a sight-reading exercise you play on your real instrument into the microphone. Draw a chord, scale, or rhythm card, then you'll get some new sheet music and a count-in, then you play. We grade every note live for pitch and timing, and if you clear the bar, the card lands. Miss, and the Choir sings back.
Six acts, eighteen foes, no retries. Bosses trash-talk you out loud in real synthesized voices, and the insults get sharper the worse you play, built from your actual grading data, so they call out the specific note you flubbed.
Six instruments. Trumpet, clarinet, alto and tenor sax, flute, French horn. Every exercise transposes into your instrument's written key, so you're reading the part you'd actually see on a stand.
Tavern duets. Two players, two parts, one phrase - recorded and played back together for the crowd.
"Gems and I." Castor and Pollux, twin AI mentors, build a daily practice set around your measured weaknesses. When you finish, they replay your own recording with the staff scrolling in time, stopping on each mistake to explain what went wrong in their own voices before you continue.
No mic? Demo mode simulates a take. No account? You play as a guest.
How we built it
Next.js 16 + React 19 + TypeScript, with Zustand for game state and MongoDB Atlas behind accounts, leaderboards, and the mentor's player file.
The heart of it is the audio path. We take the raw mic signal through Pitchy (McLeod pitch method) over Web Audio, gate each frame on level, clarity, and plausible range, then run a stabilizer so a wobbly attack doesn't read as three different notes. Grading maps each written note to a time window, skips the attack transient, and decides by pitch coverage across the window, with onset detection off a sharp tail envelope for timing. Every judgment reduces to { status, playedMidi, onsetOffsetMs }, and everything downstream - the staff colors, the taunts, the mentor's advice, the weakness profile - is built from that one honest measurement.
Gemini generates the mentor's practice plans and coaching from measured facts only (never audio), validated against a schema before anything reaches a player. ElevenLabs gives the bosses and the twins their voices, served through a signed, short-lived envelope so the browser can never turn our API key into an open text-to-speech proxy.
Notation is drawn by hand in SVG: real key signatures, beam groups, accidental spacing, and ledger lines. We needed the staff to be a live surface: cursors, hit/miss coloring, and ghost notes showing what you actually played against what was written.
Challenges we ran into
Pitch detection is not a solved problem in a browser. Getting from "there is a frequency" to "you played a C♯, and you were 90 ms late" took layers: level and clarity gates, a stabilizer, attack-skipping, coverage thresholds, onset detection on a separate fast envelope. Every one of those exists because a naive version got something wrong in testing.
Transposition bugs hide in plain sight. We store all music in concert pitch and transpose per instrument at render time, but the grader returns played pitches in a different space. Comparing them directly produced feedback that looked perfectly fine on trumpet and was an octave off on tenor sax. It only surfaced because we tested against a shift: -12 instrument specifically.
Recording and replaying a performance is all edge cases. MediaRecorder finalizes empty if its source tracks have already ended, so stopping the mic before the recorder silently produced zero-byte clips. And our clip offset is recorderStart − downbeat, which is negative because recording begins during the count-in, so adding it where we should have subtracted clamped every seek to zero and played count-in clicks instead of the passage. Both were invisible to unit tests and obvious the first time a human pressed play.
A singleton mic shared across screens. Rapid navigation could leave a queued animation frame reading through a null analyser, crashing the app.
Accomplishments that we're proud of
It actually listens. You pick up a real trumpet, play a real phrase, and the game responds to what came out of the bell. That's not a metaphor or a rhythm-game approximation.
The feedback is specific and honest. Not "78%" but "at bar 1, beat 2 you played a D instead of a C♯, a semitone above." Every word traces back to a measurement, and the mentors are explicitly forbidden from inventing anything they couldn't have measured.
The mentor closes the loop. Playing back your own recording, watching the cursor track the audio, and having it stop exactly where you went wrong is the thing we most wanted and least expected to finish.
It degrades gracefully. No microphone, no database, no API keys, and the game still runs. Every external dependency has a fallback.
Full notation from scratch, in SVG, correct across every key signature and every transposing instrument.
What we learned
Measure first, talk second. Every feature got better when we made it derive from real data rather than generate plausible text. The mentors are constrained to measured facts, and that constraint is what makes them trustworthy.
Treat the model's output as untrusted input. Gemini's plans are schema-validated and range-checked before a player sees them, and the browser can't choose what ElevenLabs says. Designing those boundaries up front saved us from a whole category of problems.
Music domain knowledge is a real engineering constraint. Transposition, enharmonic spelling, beaming, fermatas: none of it is decorative, and getting it wrong is immediately visible to any musician who looks at the staff.
Some bugs only a human can find. The clip-sync bug passed every test we could write. It took one person pressing play and saying "that's the count-in" to catch it.
What's next for SlayTheChoir
Make the buff mean something. Training rewards are currency today. They should shape your run based on what you practiced, with a pitch-focused session buying a wider pitch tolerance and a rhythm session a more forgiving timing window.
Streaks and long-term progress. The player profile already tracks weaknesses across every mode. Surfacing that ("your C♯ is 20% more reliable than last week") is the payoff we haven't built yet.
More instruments, and the rest of the orchestra. Strings and low brass need different range handling and a bass clef.
Ensemble play beyond duets. Tavern proves two players in time. Sections and full pieces are the obvious next step.
Real repertoire. Public-domain literature graded the same way, so practicing a piece you care about counts as playing the game.
Take it to a band room. The honest test is whether a director would use this for weekly sight-reading. That's the audience we built it for.
Built With
- geminiapi
- mongodb
- next.js
- pitchy
- zustand





Log in or sign up for Devpost to join the conversation.