Inspiration

We were looking for a task narrow enough that a 4B model could beat a frontier model at it, and one where the answer could be checked exactly. CAD-as-code worked for both, since a generated program either builds the right shape or it doesn't, and the difference can be measured in cubic units without asking a second model for its opinion.

The first version we tried was text-to-CAD, where the model reads a written description of a part and writes the CadQuery program that builds it. We ran the frontier models on that before spending any GPU time. Kimi K3 built 95% of our held-out parts correctly from a written spec with no worked examples, and GLM-5.3 managed 87.5%, which left very little room for a smaller model to do better.

However, engineers work from drawings, so we kept the same parts and changed only the input, re-rendering each one as a drawing sheet with front, top and right views at a single shared scale plus an isometric view for orientation. On the same benchmark Kimi K3 fell from 95% to 61%, and on parts with 13 or more faces both it and GLM-5.3 Flash dropped to 18.6%. An earlier pilot restricted to parts with 10 or more faces had already put the two of them at 22.5%.

The fact that we lost 34 points purely from changing the input format while the underlying geometry was kept identical told us the drawing itself was the hard part. Reading a drawing is a bounded visual skill, which is the sort of task we thought a specialist could be trained to do well on.

What it does

Cadabra takes a drawing sheet and the part's overall bounding box, writes a CadQuery program for it, and we then execute that program to produce a solid.

The model also checks its own work. It samples several candidate programs for one drawing, builds each of them, and re-renders every result through the same renderer that drew the input sheet. Each candidate is then scored by taking the intersection over union of its silhouette against the drawing's in each of the three orthographic views, averaging those, and multiplying the result by how closely the candidate's own bounding box matches the one stated in the prompt. That is the same IoU the grader uses, measured on pixels instead of volume, and the top-scoring candidate is the one we keep. Because the verifier only ever sees what the model saw and never looks at the reference solid, the same procedure still works on a part nobody has an answer for.

Model Correct [95% CI] Medium (7-12 faces) Complex (13+ faces) $ / 1K parts Latency p50
Cadabra 4B, best of 8 (ours) 80.7% [77-84] 61.3% 39.0% $1.19 2.3 s
Cadabra 4B, 1 sample (ours) 74.8% [71-79] 52.1% 25.4% $0.22 2.0 s
GLM-5.3 Flash (high) 66.2% [62-70] 37.0% 18.6% $1.00 3.7 s
Kimi K3 (high) 61.8% [58-66] 30.3% 18.6% $36.82 10.4 s
Qwen3-VL-4B, untuned 14.7% [11-18] 8.4% 1.7% - -

Every lane answered the same 497 held-out parts, and neither those parts nor any variant of them appeared in training. We beat both frontier models overall and on every complexity tier, and the gap widens as the parts get harder. On medium parts we reach 52.1% against 37.0% and 30.3%. Kimi K3 costs 167 times more per thousand parts while scoring 13 points lower.

On top of the benchmark we built a live 3D race, where you pick a held-out part, all three models run it at once, and each prediction is overlaid on the target solid while accuracy, latency, tokens and cost per part fill in as the answers land. Watching a model confidently build the wrong object explains the problem better than the table does.

How we built it

The data came from CAD-Coder, an open dataset pairing CadQuery programs with text specifications. We ran every reference program and rendered the resulting solid as a drawing sheet. The drawing came from the program, so the two can never disagree. That gave us 30,260 training sheets, weighted toward the medium and complex parts where the frontier models were failing. No source part appears in both the training and benchmark sets, and we dropped 1,886 training parts whose bounding box and face count matched a benchmark part. The dataset repeats simple shapes constantly, and a near-duplicate would leak straight across the split.

Training was LoRA supervised fine-tuning of Qwen3-VL-4B-Instruct on one H100 through Baseten Training, at rank 64, with adapters on the language model while the vision tower and merger stayed frozen, and loss computed on the generated code alone. A full epoch came to 1,892 steps, about 2.9 hours. Because Baseten syncs checkpoints as they're written we could benchmark intermediate ones while the job was still running, and watched the model climb from 13% untuned to 61% at step 200, 68% at step 1,000, and 73.5% by the end of the epoch. We then got four H100s and ran a second epoch with DDP, 946 steps in 55 minutes, which reached 77.0%.

We run each generated program in a sandbox, then build the solid and measure how much of its volume overlaps the reference.

$$\text{IoU} = \frac{\text{volume}(A \cap B)}{\text{volume}(A \cup B)}$$

A prediction counts as correct at an aligned IoU of 0.9 or better, taking the best score across the 24 axis-aligned rotations after centering both solids, since the dataset applies its own rotations inconsistently. Alignment doesn't soften the metric. Rotating only fixes which way a part faces. It cannot fix a wrong size, and a part built 10% too large still scores 0.75.

The verifier is separate machinery, and blind on purpose. It sees the input drawing and the stated bounding box, never the reference solid. Across 689 real frontier answers on 237 parts it told correct programs from incorrect ones with an AUC of 0.96, and asked to choose among several answers for one part it picked a correct one 98.1% of the time, against 76.3% for a random pick. It isn't specific to our model either. Kimi K3 best-of-4 through the same verifier went from 27.7% to 48.3% on the hard parts, though at $252 per thousand parts nobody would ship that.

For the frontier baselines we used Baseten Model APIs at high reasoning effort and gave each of them two worked examples. Ours gets neither, just the zero-shot prompt it trained on.

Challenges we faced

The dataset turned out to be the hardest part of the project, and we only found that out by running it. We executed all 82,659 reference programs across the dataset's splits and checked each resulting solid against the dimensions its own specification stated, which showed that 14% of the official test split contradicts itself on single-part rows where the size is checkable, against 2.3% for the hand-checked training split. Some references ignore the rotations their own text gives, and a few write files to disk while being evaluated, which is how the sandbox ended up with an import allowlist in the first place. We rebuilt the benchmark out of the clean split, grouped it by source part, deduplicated by geometry, and rendered every sheet from its own reference program.

Giving up on the original idea was the second difficulty. Our initial plan was Text-to-CAD, but the pilot killed it before it cost us any GPU time. Switching the input modality let us keep the grader, the data and the deployment path we'd already built while turning a solved problem into an open one.

Our first attempt at the render verifier scored candidates on their visible edges and didn't work, because two programs can build the same solid through different operations and end up with completely different seams. Edge matching sat near noise at 0.81 AUC on its own and pulled the combined signal down to 0.92, while silhouette agreement multiplied by bounding-box agreement got us to 0.96. The size term earned its place there, since the renderer normalizes scale and silhouettes alone can't see a part that's been built uniformly too large.

What we learned

Benchmarking the strongest models available before training anything is the habit we'd keep. Our task sounded hard and was mostly solved already, as proved in our pilot run. This changed our objective. We wanted to recover as much (and even try to beat) functionality of the stronger model on a specific task for much cheaper.

Clean evaluation was also much harder than training the model. Auditing the references, keeping source parts out of both splits, dropping requests that were never answered, and grading geometry instead of asking a model, and without it our numbers would have looked convincing while meaning very little. The 14% label error rate in the public test split makes the point on its own, since anyone reporting accuracy against that split is reporting partly on noise.

The last thing we took away is about cost. A small model doesn't have to be good at everything, only at one thing you can verify, and having a verifier that's both cheap and blind is what makes test-time sampling affordable when it mostly isn't at frontier prices. Best-of-8 on our model costs $1.19 per thousand parts, still 31 times cheaper than a single sample from Kimi K3 at $36.82.

What's next

Longer term, the interesting question is how far this pipeline can move toward real drawing-to-CAD reconstruction: taking an imperfect engineering drawing, recovering the intended dimensions and structure, generating editable geometry, and automatically checking whether the result is consistent with the source. The current system is a deliberately constrained version of that problem, but the same ingredients such as multimodal reasoning, geometric generation, deterministic verification, and cheap test-time sampling which should carry over.

Paper

Read the paper detailing our training methodology here!

Built With

Share this project:

Updates

Submission history