Inspiration

RudCo supplies metals, alloys and abrasives from 30+ suppliers, and every product page ends in "Request Quote". Buyers write RFQs however they like: "304 SS sheet 16ga 4x8" in one email, "Stainless 304, 0.060 in, 48 x 96" in the next, and both mean the same sheet. An LLM can read both, but it will also confidently guess a grade, a unit or a thickness, and in industrial supply one wrong guess means a wrong quote. We wanted an agent that is as fast as an LLM and as trustworthy as a spreadsheet.

What it does

Upload an RFQ as an email, PDF, Excel/CSV file or pasted text. For every line, the agent:

  • Matches it to one of RudCo's 61 catalog materials (steel shot sizes, aluminum ingot alloys, master alloys, abrasives).
  • Prices it against a benchmark and writes a short note on why the price is competitive.
  • Checks compliance: mill test reports, US origin, DFARS, packaging such as "25 kg bags only".
  • Asks instead of guessing. Every open question goes to the buyer in one numbered message ("1A, 2B"). The buyer replies once, only the affected lines re-run, and the quote is ready to accept.

Our rule is: code decides, the LLM copies, the buyer confirms.

How we built it

  • LangGraph agent. RFQ lines are processed in parallel; then a single interrupt sends one batched clarification, and the graph resumes from a MongoDB checkpoint.
  • LLM through the Nicrron gateway. The model only copies the buyer's exact words into a fixed schema; it never converts or infers. Questions and price notes are also LLM-written, but checked by code.
  • Deterministic resolver. Every catalog item's id is built from a closed list of allowed values, for example steel_shot|S330 or stainless_sheet|304|2B|t_0_0625|w48in|l96in. Only one function can create items, so the same product can never end up with two ids, and free text never creates new items.
  • Thickness by curated alias, not tolerance. A stated decimal $d$ matches a stocked value $t$ only when $|d - t| \le 0.0005$ in, directly or through a curated alias. So 0.060 in becomes 16 ga (flagged as an assumption), while 0.065 in becomes a question.
  • Pricing. The quote is $p = c\,(1 + m)$, where $c$ is the cheapest supplier offer that meets the buyer's requirements and $m$ is the family margin. It is compared with the benchmark $b$ as $\Delta = \frac{p - b}{b} \times 100\%$. A validator rejects any explanation containing a number that isn't among the computed fields.
  • Compliance as YAML rules. Each requirement is satisfied, unconfirmed or blocking. Unknown is always shown as unconfirmed, never yes.
  • Stack: Python 3.11, FastAPI, MongoDB, a React UI, Docker Compose, and BlockConvey PRISM tracing with one session per RFQ.

Challenges we ran into

  • The gateway's structured output. Claude 5 models reject temperature, the JSON-schema output mode was silently ignored, and list arguments arrived double-encoded as JSON strings. We switched to tool calling, Pydantic validation and a decoder at the boundary.
  • PRISM added latency. The SDK sent traces synchronously at the end of every run, so we moved delivery to a background thread that fails open.
  • Wrong answers that looked right. Live runs caught three bugs our tests had missed:
    • "40 pcs" of ingot plus the answer "lb" silently became 40 lb.
    • "MTRs required for all alloys" blocked an abrasive line. - Alloys that don't list Mg at all were offered as matching "Mg 0.5%".
  • Ambiguity is the norm. "356" could be 356.1 or A356.2; "80 grit" could be mesh or FEPA; "aluminum oxide" could be white, brown or zirconia. Each had to become a question, never a pick.
    ## Accomplishments that we're proud of
  • On a 73-line synthetic eval (emails, spreadsheets, PDF tables, government forms): 100% item match, 100% question precision and recall, 0 silent choices on ambiguous lines, and 0 invented numbers in price notes.
  • Four supplier wordings ("S330 shot 50#", "Size 330", "S-330 50 LB", "SAE S330 bag") all resolve to the same item.
  • A full human-in-the-loop flow that survives restarts: upload → one question message order.

What we learned

  • The most valuable thing an LLM can do here is less. Restricting it to copying stable.
  • "Ask, don't guess" needs structure. Numbered questions, lettered options, merged duplicates and a cap on follow-up rounds turn clarification into one tap instead of an email thread.
  • Evals must target behaviour, not just accuracy. "Zero silent choices" mattered m

What's next for RUDCO-Parsing RFQ

  • Real data: replace the synthetic suppliers and placeholder LME benchmark with Ru and have RudCo verify the curated size lists and compliance rules.
  • Recover lost questions: near-match retrieval with LLM ranking on soft signals (buyer history, application), with the buyer still confirming.
  • Harder inputs: OCR for scanned PDFs, and a review-queue UI so staff can add the solve.

Built With

  • blockconvey
  • langgraph
  • nicrronninoo
  • python
Share this project:

Updates

Submission history