Lie: Claude Code For Music Generation
A local-first, agent-first desktop music creation workstation that bridges natural language and structured composition.
Inspiration
Creating music today is caught between two frustrating extremes:
- The Pro-DAW Barrier & Fragile DAW-MCP Integrations: Traditional Digital Audio Workstations (Ableton Live, Logic Pro, Cubase) have an exceptionally steep learning curve. Novices and intuitive creators get bogged down by intricate routing, complex automation, and obscure plugin parameters before translating a single idea into sound. Recent attempts to bridge DAWs with LLMs via "DAW MCPs" or external scripting interfaces prove clunky, brittle, and exhausting to operate—forcing the LLM to micromanage GUI states and brittle plugin IDs instead of reasoning musically.
- The End-to-End Generative AI Dead End: Conversely, modern generative music models act as unpredictable "black boxes." Users are trapped in a pure "gacha pull" loop—prompting an end-to-end model and receiving an un-editable, monolithic audio block. Creators cannot iterate: you cannot tell a black-box model to "keep the groove, but rewrite only the guitar solo in bars 16–24", "switch the bassline to a syncopated walking pattern", or "add a brass counter-melody during the chorus".
Meanwhile, Large Language Models excel at symbolic reasoning, structural decomposition, and goal-driven planning. By empowering LLMs with domain knowledge bases and specialized Agent Skills, we can inject deep musical craft—harmonic theory, counterpoint, rhythmic feel, and genre grammar.
Our motivation is to build Lie as a Music Harness:
- Not a black box that replaces human artistic agency, but an engineered, deterministic foundation that maps high-level musical ideas to structured compositional data;
- An intelligent workbench that drastically lowers technical friction while unlocking the user's authentic creative and aesthetic potential;
- A system that guarantees a true Human-in-the-Loop paradigm—enforcing bounded permissions, candidate isolation, and full auditability so that music composition gains the rigor, predictability, and velocity of modern software engineering.
What it does
Lie (Agent Music Workstation) is an open, local-first, agent-first desktop workstation designed for structured, long-term music creation without DAW bloat.
- White-Box Single Source of Truth (Canonical ABC): Compositions are authored not as flat waveforms, but as human-readable, highly structured Canonical ABC notation. From this single canonical file, multi-track Standard MIDI, openDAW playback runtime states, and lossless audio exports are deterministically derived.
- Fixed Six-Role Orchestration: Provides a standardized, professional six-track ensemble—Drums, Bass, Guitar, Keyboard/Piano, Strings, and Winds—ensuring transparent arrangement hierarchy without track proliferation chaos.
- Strict UI Selection & Task Scope Bounding: The measures and tracks highlighted by the user on the timeline UI define the exact Task Scope. Natural language directs what to compose; UI selection governs where to write. The Agent cannot modify content outside this scope unless the user explicitly reviews and approves a Scope Extension.
- Candidate Transaction & Isolation Model:
- The project maintains an immutable, verified
Currentrevision alongside an isolatedCandidateworkspace. - The Agent writes exclusively to the
Candidatevia local Git worktrees without mutating the stable master state. - Users can seamlessly A/B audition between
CurrentandCandidate, and either Accept the changes to advance the project or Reject them to roll back immediately.
- The project maintains an immutable, verified
- Modular Music Style Skills: Ships with built-in, decoupled skill directories for distinct genres (Shoegaze, Midwest Emo, Metal, Post-Rock, Pop, Rap). These skills model the underlying craft—harmonic density, groove language, vocal roles, and dynamic energy curves—rather than mindlessly regurgitating surface melodies.
- Production-Ready Export: Enables one-click export to multi-track Standard MIDI and offline-rendered, verified WAV audio.
How we built it
We built Lie on a decoupled, multi-process, and strictly local architecture:
- Process Topology & Supervision:
- Electron Main (Service Supervisor): Manages process lifecycles, security boundaries, and typed IPC message routing.
- React 19 + TypeScript UI (Renderer): Features a responsive multi-track timeline, transport controls, loop selectors, and real-time playback powered by openDAW and SpessaSynth SoundFont rendering.
- Music Core Utility Process: An independent Node.js backend managing project models, Git worktrees, ABC compilation, and deterministic state machines.
- Agent Runtime Process: Built on the Strands framework, hosting agentic workflows, LLM providers, and an MCP client.
- MCP (Model Context Protocol) Security Surface:
- The Agent has zero direct access to the filesystem, shell, or Git. All modifications are governed by 9 domain-specific MCP tools (e.g.,
getScopedComposition,replaceScopedMusic,resizeComposition,updateMusicalProperties,submitGenerationPlan).
- The Agent has zero direct access to the filesystem, shell, or Git. All modifications are governed by 9 domain-specific MCP tools (e.g.,
- Deterministic Music Validation Pipeline:
- Custom ABC parser and validation engine enforcing strict barline integrity, tie-chain continuity, MIDI velocity standards (1–127), and singular Tick-0 tempo/meter authorities.
- Structured Long-Form Arrangement:
- Long tracks are constructed by decoupling structural timeline sizing (
resizeComposition) from editing permissions, allowing the Agent to iteratively plan and compose across sequential sub-scopes (targetScope).
- Long tracks are constructed by decoupling structural timeline sizing (
Challenges we ran into
- Symbolic Music Precision & Hallucination Mitigation: LLMs frequently produce misaligned ticks, unequal bar lengths across tracks (
TRACK_LENGTH_MISMATCH), or invalid accidental notations. We solved this by implementing an actionable validation engine that aggregates syntax errors and returns concise, deterministic hints (e.g., exact tick discrepancies per voice), closing the loop via an automated self-repair cycle. - Decoupling Composition Structure from Authorization: In early designs, expanding the timeline length was conflated with granting edit permissions. We completely decoupled
resizeComposition(which modifies the project's measure length) fromrequestScopeExtension(which only expands write authorization), creating an idempotent, fail-closed foundation for arbitrary-length tracks. - Asynchronous Human-in-the-Loop Transactions: Handling user cancellation, project switching, and unexpected agent crashes required a robust transaction model. We introduced a non-expiring
Operationstate model and automatic crash reconciliation, ensuring that any interrupted run safely rolls back to the authoritative Git base checkpoint without corrupting project state.
Accomplishments that we're proud of
- Pioneered White-Box AI Composition: Moved beyond "black-box toy audio generators" and clunky DAW automations to deliver an open, fully editable, structured workstation where every note, velocity, and arrangement decision remains in the artist's hands.
- Engineered a Rock-Solid Music Harness: Implemented software-engineering-grade discipline (Git worktrees, MCP authorization fences, candidate staging, and A/B comparison) into a creative music workstation.
- Built Expressive, Craft-Centric Style Skills: Created domain skills that analyze genre DNA—rhythmic syncopation, harmonic colors, and arrangement density—teaching LLMs to think like seasoned music producers.
- 100% Local-First Privacy & Zero Cloud Lock-In: The entire core, version history, SoundFont synthesizers, and workspace data reside locally on the user's machine, keeping the artist's creative IP completely private and secure.
What's next for Lie
- Adapting Next-Gen Music Models (e.g., yue2): Actively integrating and benchmarking next-generation open models like yue2, exploring a hybrid paradigm that pairs Lie's structured symbolic harness with cutting-edge neural acoustic rendering to dramatically optimize generation costs and acoustic expressiveness.
- LLM-Programmable Timbre & Custom DSP Effects System: Building an intelligent, programmable timbre engine where the LLM can directly author, configure, and modulate software audio effects (reverbs, distortions, filters, modulation chains) in response to artistic descriptions—allowing artists to generate bespoke DSP effects and curate their own signature sound libraries.
- Hybrid Graphical & Symbolic Editing: Expanding the React timeline to include native Piano Roll, Guitar Tablature, and Automation Curves, enabling seamless transitions between natural language guidance and manual note-level tweaking.
- Open Community Skill & SoundFont Ecosystem: Standardizing our Agent Skill specifications to allow musicians and the open-source community to contribute custom genre guides, progression templates, and custom SoundFont (SF2/SFZ) instrument collections.
Log in or sign up for Devpost to join the conversation.