Inspiration
A voice lesson is eighty dollars and forty-five minutes. Most of that time is the same loop: listen to a take, name what went wrong, pick an exercise, listen again, try to remember last Tuesday.
Singers who practice alone do not get that loop. They get a karaoke app that prints a number, or a chatbot that prints a paragraph about appoggio. Neither watches this take. Neither assigns a drill the app can grade. Neither comes back tomorrow. Beginners also cannot hear “support” or “placement” — they need to see the instrument, be told what to sing next, and be heard again.
That loop is repetitive skilled work. It is exactly what this hackathon asked for.
LetsSingAI gives the unpaid half of a vocal lesson to a Strands agent on Amazon Bedrock. The primary user is a singer doing a craft — or a teacher reviewing a session log — not someone paying a bill. That is why this is Professional Agents, not Everyday. The singer’s only job is to sing. The agent owns listen → diagnose → drill → pass/retry → remember.
Karaoke scores you. Chatbots talk about singing. This agent runs the lesson.
What it does
Try it: https://letssingai-production.up.railway.app/ · Chrome · allow the mic.
A session is one sitting, not a chat.
- Pick a phrase. Public-domain Twinkle, Ode to Joy, or Amazing Grace. Hear it. Optionally play a real demo voice through the same analysis as the mic.
- See the instrument. An orbitable 3D anatomy view — lungs, folds, tract, placement — driven by your voice. Before you sing, Show correct technique walks the target motion. While you sing, a translucent ghost shows the correct shape next to you, and the part that’s off glows. This is an educational acoustic model (pitch → fold tension, formants → tract shape, loudness stability → breath). It is not medical imaging, and the UI says so.
- Sing. Pitch, formants, harmonics, and a breath-stability proxy are extracted in the browser. There is no LLM on this path. A pitch ribbon colors your trail against the melody (teal dot = the note you should be on). Harmonics show your H1–H8 against the reference timbre.
- Stop. The app does not upload audio. It sends a compact JSON summary to a Strands orchestrator. Raw audio never leaves the browser.
- The agent conducts.
vocal_coachon Amazon Bedrock:recall_singer— last session, last focusdiagnose_take— the issues, and which body parts moved wrongplan_practice— exactly three knowledge-base drills with numeric pass criteria the app actually scorescompare_progress/remember_session— so next time starts from last time- for each drill,
review_drill— pass, retry, or skip - a recap, plus a teacher-log download
- You only sing the drills. Hiss, siren, slow-note matching, ng-to-ah, yawn-sigh — whatever the plan named. You do not pick the exercise. The agent does.
The UI shows the Strands tool loop as chips while it happens. Lesson steps stay on screen: Take → Drill 1 → Drill 2 → Drill 3 → Recap. A “last time” strip shows prior cents, breath, and focus.
If Bedrock is down, the same tools still run on a heuristic path and the badge reads offline heuristic. We do not fake a Bedrock session.
Pitch is scored the way a listener hears it, not the way a studio tuner does. The centre of a held note is compared to the written melody; scoops, vibrato, and a slightly late landing are forgiven. A perfectly centred wrong scale degree still fails. That is the difference between “you were 8¢ sharp of the nearest piano key” and “you sang the wrong note.”
Four techniques live in the knowledge base, each with three runnable drills and pass/fail metrics: breath support, pitch accuracy, forward resonance, open throat. The agent is not allowed to invent a drill id.
How we built it
Two clocks. The instrument has to move at 60 fps. The coach can think for a couple of seconds. Mixing those is how vocal-AI demos become sluggish chat overlays. We split them.
The instrument (no LLM)
React 19, React Three Fiber, three.js, pitchy, Web Audio, Zustand, Vite.
- Live DSP: f0, LPC formants, H1–H8, breath-stability proxy
- 3D anatomy in three modes from one pose model: live, you vs target ghost, correct-technique demo
- Melody-relative cents (not nearest-semitone)
- Phrase picker, play-along, demo voice, teacher-log export
The conductor (Strands on Bedrock)
FastAPI streams POST /coach over SSE (event=take|drill|recap). Events: focus, diagnosis, plan, verdict, tool, token, done.
A Strands agent named vocal_coach is the session conductor. Tool choice is model-driven.
| Job | Tool | Constraint |
|---|---|---|
| Remember who this is | recall_singer / remember_session / remember_note |
MemoryStore protocol |
| Name the fault | analyze_take → diagnose_take |
Grounded in the JSON; no invented metrics |
| Assign homework | plan_practice / get_exercises |
KB ids only, exactly 3 |
| Grade the drill | score_drill / review_drill |
Pass criteria from the KB |
| Track trend | compare_progress |
Last session vs this one |
On Claude Sonnet, diagnose_take, plan_practice, and review_drill nest as specialist agents-as-tools (narrow prompt, small tool set). On Nova Lite (us.amazon.nova-lite-v1:0, the default so anyone can run without an Anthropic form), the same names are called directly. A ToolSinkHook streams each tool start/done into the UI chips.
GET /health → { coach: "bedrock-agent"|"heuristic", memory: "agentcore"|"local", model: "…" }.
AWS
- Amazon Bedrock — Nova Lite by default; Claude optional
- Bedrock AgentCore Runtime —
agent/Dockerfile,agent/runtime.py(BedrockAgentCoreApp) - AgentCore Memory when
AGENTCORE_MEMORY_IDis set; otherwise local JSON behind the same MemoryStore protocol
The public demo is one HTTPS origin (UI + /coach). AgentCore is packaged so the same conductor can run on AWS. Tests (pytest in agent/) cover KB-id constraints, scoring, and the coach contract without AWS.
Architecture diagram: docs/architecture.svg
Challenges we ran into
Nearest-semitone cents lie. Score every audio frame like a strobe tuner and decent singing looks “off”: scoops into the note, vibrato of ±30–50¢, a landing a beat late. Listeners do not hear that way — they judge the centre of a held note against the melody. We rebuilt evaluation around that (ignore attack/release, merge syllable-notes, ~50¢ in-tune window, 100¢ = wrong note). Then we had to make sure a perfectly centred wrong scale degree still fails, or the agent would congratulate you for singing the wrong song in tune.
The agent must not invent homework. An LLM will happily assign “do lip trills on a 9/8 pattern in G♭.” The app cannot grade that. Tools only return exercises from the JSON knowledge base; plan_practice returns exactly three; the UI only runs those ids. Tests lock this.
The 3D model cannot wait on Bedrock. Keeping live visuals off the LLM was a product constraint, not a shortcut. Contract: JSON in, streamed markdown + a drill plan out. The expensive model sees a summary, never a wav.
Showing the agent without turning it into theatre. Judges (and singers) need to see the loop, or this looks like “a chatbot with graphics.” Tool chips are the difference. The honest fallback was harder than a spinner: if Bedrock is down, the badge says so.
Anatomy without lying. Pitch, formants, and loudness are real. A glowing larynx is an inference. We labeled it on the canvas, in the system prompt, and here.
Accomplishments that we're proud of
The agent runs the lesson. It does not chat about singing.
A complete take → three scored drills → pass/retry → recap loop, with memory, a teacher log, and a 3D instrument that moves while you sing — no model in that path.
A working public demo whose /health reads "coach": "bedrock-agent", plus AgentCore packaging, plus a fallback that does not pretend.
A public Apache-2.0 repo that reaches a Strands · Bedrock badge on Nova Lite without an Anthropic form.
What we learned
A wrong note sung in tune should fail. That one sentence changed the product from a tuner with a speech balloon into a coach.
Judges and singers both need to see the agent loop. Tool chips, lesson steps, and an honest badge do more than a paragraph of architecture.
Nova Lite is enough for a real Strands conductor. Nested specialists are a Claude bonus, not a requirement — same tool names, same lesson.
Pedagogy belongs in the tools. If pass criteria live in a knowledge base the app can score, the model cannot wander. If they live only in the prompt, it will.
What's next for LetsSingAI
A guided metronome for the slow-note drills (the agent already assigns them; the app should click the pulse). More songs through the existing librosa analysis pipeline. AgentCore Memory strategies — preferences and teacher notes, not just session events — when a memory id is set. More phrases, same conductor.
The loop does not change: you sing, the agent runs the lesson.
Built With
- amazon-bedrock
- amazon-nova
- bedrock-agentcore
- fastapi
- librosa
- pitchy
- python
- react
- react-three-fiber
- sse
- strands-agents
- three.js
- typescript
- vite
- web-audio
- zustand
Log in or sign up for Devpost to join the conversation.