Inspiration
Every company runs on tribal knowledge: who has to sign off before a file is touched, how a service moves from staging to prod, which migrations need to be phased, which "simple" changes are secretly dangerous. It was never written down. It built up over years of chats, tickets and code reviews.
Companies are now dropping coding agents into these codebases, and the agents have none of that knowledge. Our motivating example:
A new hire asks an agent to "make the primary button blue." The agent was opened inside
packages/ui/button/, sees three files, and makes a one-line change. It has no idea it's sitting inside a monorepo, and thatButton.tsxis a shared component imported by every app in the company. One line of CSS changes thousands of screens.
The agent isn't careless. It doesn't know what it doesn't know.
The industry's answer is retrieval: dump PRs, tickets and chat into a vector database and let the agent search when it's unsure. But retrieval is only consulted when the model decides it's confused, and a confident model on a task that looks trivial never asks. When it does ask ("policy for changing buttons"), a decade of unstructured history returns hundreds of loosely related documents, and none of them say "this component is shared."
So we asked: what if the agent didn't have to look it up at all? What if the knowledge lived in the weights?
What it does
weave fine-tunes a coding agent on how your company actually works.
- Connect a Discord bot and a GitHub repo. That's the whole setup.
- We rebuild your history as a series of tasks: every Discord message, Jira change and GitHub diff that went into landing one ticket.
- A supervisor model replays each task inside our own harness, the way a senior engineer with full context would have solved it, and we record every tool call.
- We fine-tune on those replays and host the model behind an OpenAI-compatible API, served through our desktop harness.
Given the same button ticket, the fine-tuned agent recognizes packages/ui as a shared library and adds a checkout-only variant behind a feature flag. It doesn't search, because it already knows.
How we built it
The pipeline is a medallion architecture built end to end on Databricks, with Lakebase (Postgres) as the single source of truth for every stage.
| Stage | What happens | Databricks |
|---|---|---|
| Bronze | Raw GitHub commits/PRs/events, Jira issues and changelogs, and Discord messages, landed as-is | Remote Extensions → Lakebase |
| Silver | Drop malformed JSON, missing fields, broken timestamps and bot messages; link commits ↔ tickets ↔ threads. Failed rows go to silver.quarantine rather than being deleted |
Notebooks |
| Gold | A model proposes task windows; static code verifies every boundary; each window is sliced into chunks | AI Model Gateway + Notebooks |
| Platinum | Each window becomes a sandbox (TASK.md, CHAT.md, EVENTS.md, COMMITS.md, REPO/); an agent replays it with the exact 10 tools our harness ships |
AI Model Gateway |
| Weights | LoRA fine-tune on the replay transcripts | Databricks GPU job |
A scheduled trigger reruns the whole pipeline every 12–24 hours so the model keeps up with the codebase.
The model proposes, the code decides. Window detection is the hardest stage, because nobody ever labelled when a task started or ended. The model has to cite its evidence, and a window is kept only if every cited row actually falls inside its span:
$$ \text{keep}(w) \iff \forall\, r \in \text{evidence}(w):\; t_{\text{start}}(w) \le t(r) \le t_{\text{end}}(w) $$
Surviving windows are scored against a deterministic, rules-based baseline, and if the model doesn't beat the baseline, the baseline wins.
Training on the trajectory, not the answer. Each training example is one full conversation: system prompt, task, and every ordered tool call. Loss is applied to assistant turns only; tool outputs are masked so the model never learns to hallucinate its own tool results:
$$ \mathcal{L} = -\sum_{t=1}^{T} m_t \log p_\theta(x_t \mid x_{<t}), \qquad m_t = \begin{cases} 1 & x_t \in \text{assistant turn} \ 0 & \text{otherwise} \end{cases} $$
The harness. An Electron + Next.js desktop app with the same ten tools used during replay (list_dir, read_file, write_file, edit_file, search, run_command, make_dir, environment, diff, list_changes), a realpath-enforced workspace jail, crash-safe undo snapshots, and an ordered transcript. Because the model is trained on exactly these tool schemas, it performs best inside exactly this tool.
The platform. A FastAPI service with a dashboard, an OpenAI-compatible /v1/chat/completions endpoint, and GitHub and Discord ingest.
By the numbers (from our live run)
- 32 task windows found and verified, sliced into 810 chunk rows
- 32 replays captured, 1,167 tool calls
- 18 validated training conversations, 131,481 tokens
- 37 rows sitting in quarantine instead of silently deleted
- 119 tests passing across the notebooks, app and ingest
Challenges we ran into
- Finding task boundaries without labels. Letting a model write directly to gold produced windows citing messages that didn't exist. Splitting the job into model proposes → code verifies → baseline must be beaten made the stage trustworthy.
- Three Discord export formats. Our bot wrote
message_idwhere the pipeline expectedid, so every message was ingesting with a null key. We wrote one parser for all three shapes and tested it against the bot's real exports, not fixtures. - Dataset validation. Before any GPU sees the data, we check that every tool call is answered, no results are orphaned, no conversation ends on a tool message, and no window appears in both train and val. That gate caught 696 orphaned tool results caused by a single wrong key name.
- Compute. The LoRA trainer is written and debugged, but on a 16 GB machine the 4B model loads, plans 60 steps, then stalls. Moving training onto a Databricks GPU job is the next step.
What we learned
- Retrieval fails one step before the search. The failure isn't a bad index; it's the model deciding the question wasn't worth asking. You can't fix that with better embeddings.
- The trajectory is the product. The final diff tells you what changed. The sequence of reads, searches and decisions tells you why, and that's what we train on.
- Never let a model write to the source of truth unsupervised. Every model output in the pipeline is a proposal checked by deterministic code.
- Lakebase as the single source of truth made every stage auditable: any gold window can be traced back to the exact bronze rows it came from.
What's next
- Complete the fine-tuning run on Databricks GPU clusters and benchmark it against retrieval-augmented agents on held-out tasks
- Run the pipeline on a real multi-year corpus instead of our synthetic one
- Add Slack and Confluence connectors
- Close the loop: every task a customer runs through the harness becomes new training data
Built With
- ai-model-gateway
- codemirror
- databricks
- databricks-notebooks
- databricks-workflows
- electron
- fastapi
- github-api
- jira-api
- lakebase
- lora
- mlx
- next.js
- node.js
- openrouter
- postgresql
- python
- react
- sqlite
- typescript
- unsloth
Log in or sign up for Devpost to join the conversation.