Inspiration

Kamion (YC), Türkiye's largest freight platform, gave us the brief: value a used semi-tractor from a seller's photos, knowing that sellers sometimes lie. The obvious build is one vision-model call that returns a number. The brief says out loud that this is the thing that will not win — and it's right. A model that prices and describes in the same breath has no way to be checked, and it never says "I can't tell."

So we built the ordering instead of the shortcut.

What it does

You drop in photos of a truck. You get back a price range, a condition report where every finding names the photo it came from, and — when the photos don't support a price — a refusal that says what shot to take next.

Four of our eight rehearsed demo cases produce no price at all, each for a different reason:

  • close-ups only, never the whole truck → condition notes, no price, asks for the shot
  • three different trucks in one listing → names what differs, refuses to average them
  • a motorcycle → refused before a single vision token is spent
  • unreadably dark and blurry → refused on measured capture quality

Knowing its limits was its own line in the rubric, so we made the refusals tested cases with expected outcomes rather than something to improvise on stage.

How we built it

photos ─▶ Gate ─▶ Heads ─▶ Evidence ─▶ Reconcile ─▶ Price ─▶ Report

The ordering is the argument:

  1. Gate — YOLOv8n truck detection, Laplacian-variance blur and exposure checks, CLIP zero-shot view tagging. Cheap and deterministic, so it runs first and can stop the pipeline before a token is spent. Truck detection is set-level, never per-photo: a close-up of a tire is still a photo of the truck.
  2. Perception heads — three linear heads on the CLIP embedding the gate already computed, trained on 3,729 exactly-paired clean/degraded images where the applied corruption is known. Every fold grouped by vehicle.
  3. Evidence — one structured call, all photos at once, into a fixed schema over a closed enum of 33 heavy-vehicle components. Every issue must cite a photo_id that was actually sent; ones that can't are dropped rather than trusted. The vocabulary is the point — steer vs. drive tires, fifth wheel plate scoring, frame rail corrosion, AdBlue tank, DPF. A generic dent/scratch taxonomy is the tell that a car tool was pointed at a truck.
  4. Reconciliation — the vision model writes a confidence about its own claim and that number multiplies straight into the price. Nothing else was in a position to disagree with it. Now a fine-detail claim resting on a frame the degradation head scores past its cutoff gets downgraded, and the downgrade is printed. Nothing is deleted — changing your mind is allowed, doing it quietly is not.
  5. Price — a hedonic ridge on harvested asking prices, blended by inverse variance with a second route: published new price × a fitted retention curve, which needs no same-brand comparable to exist. The model never sees the photos; the vision model never sees a price.

We also built the dataset: 200 vehicles, 3,729 originals plus 3,729 degraded twins, harvested via Firecrawl from Ford Trucks Türkiye's and Daimler's OEM used networks and an independent-dealer marketplace. Every original was visually checked — tiled into 264 contact sheets and reviewed, not sampled. No image was ever removed for being badly shot; blur, mud and bad framing are the point.

On tooling: 54FA26 allowed AI tools and we used them all the way down. The code was written by agents — Claude Code primarily, with Codex and Cursor Pro alongside. Our work was directing them: what to build, what to measure, and what to throw away when the measurement came back negative.

What we measured

Judges ask which numbers are real. These are:

Quantity Value
80% price band coverage 80.3% over 958 held-out evaluations
Point accuracy R² 0.84, median error 4.2%
Gate false-refusal on real trucks 1 of 200 vehicles
Degradation head AUC 0.985, severity R² 0.695
Retention curve R² 0.936, median error 3.1%
Anchor on an unseen make band 1.85× → 1.00×, error 16.3% → 3.7%

Three quantities are stated assumptions rather than measurements, and the report says so at the point it uses them.

Challenges we ran into

A 2021 F-MAX priced at ₺118,000,000. Moving the price model's target to native currency left the estimate still multiplying by the TRY rate — and every relative assertion in the test suite sailed past a 48× error, because both numbers were wrong by the same factor. There is now an absolute magnitude test.

The 80% band covered 70%. It was being built from in-sample residuals, which are far tighter than real prediction error after a ridge fit on ~18 groups.

Folds that scored memorisation. 87 of 155 priced listings collapse into 19 identical-spec groups — one dealer lists the same 2022 F-MAX twenty-six times. Grouping folds by spec fixed it, the same reason image splits are grouped by vehicle.

Measuring the wrong thing. Unseen-brand widening came out at 1.05× when measured by dropping brand columns — meaningless on a corpus that is 93% Ford. Leave-one-brand-out across markets then reported 442% median error for Ford, because holding Ford out removes Türkiye entirely: it was measuring the border, not the brand. Done within a market, it says 1.85×.

Accomplishments that we're proud of

The measurement that killed a feature. The obvious next move was letting photos predict the part of price that age and kilometres can't explain. We ran it as a nested, vehicle-grouped probe against a permuted-target null first — six pooling variants, all negative, while the same pipeline passed its positive controls. The blocker is the target: Turkish asking prices are 23 distinct values across 84 listings. That's a dealer's price table, not a market. So the condition weights stay hand-set and stay labelled an assumption, and one script saved 155 vision calls on a fit with nothing to learn.

Also: 52 offline tests that run in half a second with no API calls, and any appraisal can be frozen into one self-contained HTML file — stylesheet, script and every photo inlined — that opens with no server and no network. Unlike a screen recording, you can still click through it when a judge asks a question.

What we learned

Refusing and re-asking are not the same answer, and neither is a wrong number. Most of the engineering went into making the system able to say which of the three it's giving you — and into measuring things early enough that a negative result was cheap instead of a wasted weekend.

What's next for KamionVision

A corpus with real transaction prices instead of asking prices. That single change unblocks the learned condition model, which is the thing we proved this data cannot support.

Built With

Share this project:

Updates