Leadsheet
Inspiration
The inspiration for Leadsheet came from recognizing a fundamental gap in how music and software interact. Professional music creation has traditionally required either:
- Expensive digital audio workstations (DAWs) like Ableton, Logic Pro, or Cubase — powerful but steep learning curve and hundreds of dollars
- Cloud-based music services with recurring subscriptions and locked-in vendor ecosystems
- Deep music theory knowledge to use traditional notation software or music libraries
Meanwhile, the rise of natural language interfaces meant users could now describe what they wanted in plain English: "Create a happy piano song" or "Write an upbeat drum pattern."
The core insight: What if music generation could be as simple as typing a description, and the output could be a readable, editable text file?
This led to the realization that:
- Musicians and software engineers think differently about music
- A text-based format could bridge that gap
- Version control and collaboration become possible if music is stored as text, not binary audio
- Local processing means no subscriptions, no data privacy concerns, and full offline capability
- Integration with common LLM-based tools could democratize music creation
Leadsheet was born to answer: Can we make composing music as accessible as writing code?
What It Does
Leadsheet is a complete music composition toolkit that translates natural language descriptions into playable, editable music files and audio.
The Three Core Flows
1. Describe → Compose → Play
"Write me a calm ambient piece"
↓
.leadsheet file created
↓
Validated, compiled, rendered
↓
Playable MP3/WAV ready to hear
2. Edit & Iterate
- User edits the
.leadsheettext file (change tempo, swap instruments, modify melody) - Recompose to hear the changes
- Cycle repeats until satisfied
- Full version control history preserved
3. Export & Share
- Audio files are tagged with metadata (MP3) or saved as WAV
.leadsheetsource files can be committed to git- Others can read, understand, and modify the source
What Users Can Create
- Chord progressions with specified instruments and voicings
- Bass lines with root-fifth patterns or walking bass
- Melodies with precise note sequences, durations, and dynamics
- Drum patterns using standard drum kit notation
- Complete songs by combining all of the above with sections and reuse
- Background tracks for presentations, videos, or creative projects
Key Capabilities
- 161+ chord types (major, minor, suspended, extended, altered, etc.)
- 128 General MIDI instruments (piano, guitar, strings, brass, woodwinds, etc.)
- Multiple drum kits with percussive sounds
- Reusable patterns (define once, use everywhere with transforms)
- Sections for song structure (intro, verse, chorus, bridge, outro)
- Transposition and dynamics control for artistic variation
- Cross-platform audio rendering with automatic fallback selection
How We Built It
Foundation: We started with audio. Built a MIDI compiler, integrated FluidSynth for synthesis, wired up FFmpeg for MP3 encoding. The pipeline worked—users could hear their music.
The Format Problem: But there was a wall: musicpy's notation was verbose. A simple two-track piano piece needed 355 tokens in JSON. For LLMs paying per token with context limits, that's a dealbreaker.
So we tested eight formats against a real corpus: lo-fi beats, rock songs, metal epics. Measured tokens obsessively. The winner—a line-oriented structure with short keywords—crushed the baseline. 91% reduction on simple pieces. 2,327 tokens down to 348 on complex metal compositions. Average 82.5% corpus-wide. That's not optimization; that's a new economic reality for LLM composition.
The insight: stop designing for humans or musicpy. Design for what LLMs can predict reliably. Short field names, explicit values, no JSON boilerplate. The DSL became clean enough for humans and sparse enough for AI.
Integration: We built a single MCP server that works across Claude Code, ChatGPT, Codex, and Gemini—each with different config paths. Smart client detection in setup made it seamless.
Audio Fallback: Users without FluidSynth and FFmpeg were stuck with unplayable MIDI. Instead of failing, we built a cascade: FluidSynth + FFmpeg → MP3. FluidSynth alone → WAV. TinySoundFont (optional Python package) → in-process WAV. None of the above → MIDI + clear upgrade instructions. Every user gets something playable; we made upgrade paths obvious.
Challenges We Ran Into
The DSL Design Puzzle
When we started designing the format, we realized it had to walk a tightrope. It needed to be readable enough for humans who knew nothing about music theory, yet token-efficient enough for LLMs to generate reliably without blowing through their context limits. It also needed to be unambiguous—no hidden edge cases where an LLM might generate something syntactically correct but musically nonsensical.
We tested eight different approaches. Some were too terse and unreadable. Some were too verbose and wasteful. We played with different notations: Am7 vs A-7, C5+E5+G5 vs C5 E5 G5, x8 vs repeat=8. Every choice mattered because it directly affected how many tokens an LLM would need to spend generating a piece.
The format that won was the one that eliminated all JSON boilerplate, used short explicit field names, and grouped related concepts per line. It scored a 91% token reduction on simple pieces and maintained an 82.5% average reduction across the full corpus. More importantly, it was something LLMs could predict and complete reliably because the structure was repetitive and obvious.
The Audio Dependency Problem
We shipped the tool, and users started reporting: "Where's my audio file?" It turned out that high-quality audio rendering requires external binaries—FluidSynth for synthesis and FFmpeg for encoding. Installing these is an OS-specific nightmare: Homebrew on macOS, apt-get on Linux, Chocolatey on Windows. Many users didn't have admin access or the patience to debug installation instructions for tools they'd never heard of.
The ugly outcome was that users without these binaries were stuck with MIDI files, which are unplayable on most consumer systems. Support tickets piled up. We faced a choice: either require the binaries and lose half our users, or find a way to make it work anyway.
We built a fallback chain. If both FluidSynth and FFmpeg were available, users got a high-quality MP3. If only FluidSynth was available, they got WAV. If neither was available but they'd installed the optional TinySoundFont package, they got WAV from an in-process renderer. As a last resort, they got MIDI plus explicit instructions on how to upgrade.
The key was that the tool never failed silently. It always produced something, always told the user what they got, and always explained how to get something better. This turned a support burden into a learning opportunity.
The Multi-Client Integration Maze
We built an MCP server. Perfect. Then we realized that MCP doesn't actually standardize how clients register servers or where they store configuration. Claude Code uses one path. Codex uses another. Gemini CLI uses yet another. ChatGPT doesn't auto-discover at all. The Skill documentation goes to different directories depending on the client.
We had to reverse-engineer each client's configuration format and build a smart setup wizard that could detect which clients were installed, register the server with each one, put the Skill in the right place, and then report back what worked and what didn't. This is invisible work—nobody notices it until it breaks. But it's the difference between "install and it just works" and "install and spend an hour debugging obscure config files."
The Music Theory Validation Gap
Early on, we discovered that users could write .leadsheet files that parsed perfectly but sounded terrible. They'd write chord symbols that musicpy interpreted differently than they expected. Or they'd voice a chord in a way that violated music theory. Or they'd mix enharmonic equivalents awkwardly.
A file that parses but generates wrong music is worse than a parse error—it undermines confidence in the tool. We added a validation layer that cross-checked every chord against musicpy's music-theory constraints. If something was ambiguous or invalid, we reported it immediately with a clear error pointing to the exact line.
The Installation Friction Problem
When users ran pip install leadsheet, they got the Python package. But they didn't get the MCP server registered, the Skill installed, or their soundfont cached. Worse, there was nothing telling them they needed to run leadsheet setup.
This is a fundamental problem with pip—there's no reliable post-install hook. So we made the setup command idempotent, made it very fast (under a second), and embedded it into the recommended install flow: uv tool install leadsheet && leadsheet setup. It's still an extra step, but it's unavoidable without fixing Python packaging infrastructure.
The Silent Track Drift Problem
In a multi-track composition, all tracks need to stay synchronized. A melody intended for 8 bars needs to actually be 8 bars. But users often miscounted—a melody that looked like 8 bars might actually be 7.5 bars of notes. When layered over an exactly 8-bar chord progression, the timing misalignment creates either sudden silence or overlapping playback.
The insidious part was that this was a silent failure. The music would render and play, but it would sound "off" in a way users couldn't quite diagnose. We fixed it by making validate always return actual track lengths and adding optional bars=<n> assertions. Now when tracks drift, users get an explicit error message pointing to the exact problem instead of a mysterious audio artifact.
Accomplishments We're Proud of
- End-to-end. Natural language to playable audio. Not a toy—production-ready composition.
- A DSL that bridges humans and LLMs. Readable syntax, token-efficient generation, composable semantics, version-controllable. Rare to nail both audiences.
- No subscriptions, no vendor lock-in. Local rendering with graceful fallback. FluidSynth + FFmpeg if you have them; TinySoundFont if you don't; MIDI + upgrade path if you have nothing.
- One codebase, four clients. Claude Code, ChatGPT, Codex, Gemini—all speaking the same MCP protocol, all reaching the same feature set.
- 161+ chord types, validated. Music theory isn't a suggestion. Every chord cross-checked against musicpy's constraints.
- Clear failure modes. When something goes wrong, users know exactly what and why. No silent errors.
What We Learned
- Text, not binary. MIDI and MP3 are for listening. Text is for understanding, version control, and LLM collaboration. A text format opened doors traditional music software couldn't.
- Graceful fallback beats hard requirements. "Users need FluidSynth" became a support problem. "Users don't need FluidSynth" became a feature. Degradation is design.
- Music theory is mandatory. Syntax validation isn't enough. A grammatically correct
.leadsheetwith invalid chords is worse than a parse error—it undermines trust. - Clear errors matter more than you think. The difference between "I can fix this" and "I give up" lives in the error message.
- Examples teach better than specs. A curated corpus of real
.leadsheetfiles teaches faster than 50 pages of grammar. Humans and LLMs both learn from patterns. - Choose your license deliberately. AGPL-3.0 means community benefits when someone builds on Leadsheet. That choice shapes the project's entire ethos.
What's Next for Leadsheet
Two parallel tracks: Product (tool improvements) and Research (local trained-model inference).
Product roadmap
- v0.3.x: Treesitter Grammar, VSCode syntax highlighting, MusicXML/PDF score export, expanded genre docs & examples
- v0.4.x: Non-4/4 time signatures & compound meter, microtonal/alternative tuning, smarter MIDI voice leading
- v0.5+: Real-time web editor, community leadsheet library, counterpoint/harmonic analysis, DAW export (Ableton, Logic), Discord community & tutorials
Research track
Train a local, open-source model to generate .leadsheet from natural language, removing dependency on external LLMs entirely.
- Build a MIDI-to-leadsheet pipeline and reverse-engineer 1,000+ symphonies into the format
- Pair each with natural-language descriptions; curate for quality and DSL gaps
- Fine-tune an open model (e.g., Qwen) for local CPU/GPU inference
Two-tier architecture
- Local: CLI + MCP server, zero telemetry, fully offline, AGPL-3.0
- Remote: hosted MCP endpoint with optional anonymous telemetry (prompt + output + validation, PII-stripped) to feed the training pipeline
Remote usage generates training data → improves the model → improves local inference → attracts more users → more data. Product and Research reinforce each other.
Vision
Short-term: a text-based bridge between natural language and playable music, usable by anyone regardless of music theory background.
Long-term: fully offline natural-language composition, powered by a model trained on real (prompt, .leadsheet) pairs — no APIs, no subscriptions.
Principles: open-source (AGPL-3.0), local-first, text-based, accessible to non-musicians, a bridge rather than a DAW replacement.
We're building toward a future where composing music is as accessible as writing text.
Project Status: Actively developed and maintained
Community: Open to contributions and feedback
License: AGPL-3.0-or-later
Built With
- ai
- chatgpt
- cloudflare
- mcp
- python
- react
- skill
Log in or sign up for Devpost to join the conversation.