Inspiration
RudCo supplies metals, alloys and abrasives from 30+ suppliers, and every product page ends in "Request Quote". Buyers write RFQs however they like: "304 SS sheet 16ga 4x8" in one email, "Stainless 304, 0.060 in, 48 x 96" in the next, and both mean the same sheet. An LLM can read both, but it will also confidently guess a grade, a unit or a thickness, and in industrial supply one wrong guess means a wrong quote. We wanted an agent that is as fast as an LLM and as trustworthy as a spreadsheet.
What it does
Upload an RFQ as an email, PDF, Excel/CSV file or pasted text. For every line, the agent:
- Matches it to one of RudCo's 61 catalog materials (steel shot sizes, aluminum ingot alloys, master alloys, abrasives).
- Prices it against a benchmark and writes a short note on why the price is competitive.
- Checks compliance: mill test reports, US origin, DFARS, packaging such as "25 kg bags only".
- Asks instead of guessing. Every open question goes to the buyer in one numbered message ("1A, 2B"). The buyer replies once, only the affected lines re-run, and the quote is ready to accept.
Our rule is: code decides, the LLM copies, the buyer confirms.
How we built it
- LangGraph agent. RFQ lines are processed in parallel; then a single interrupt sends one batched clarification, and the graph resumes from a MongoDB checkpoint.
- LLM through the Nicrron gateway. The model only copies the buyer's exact words into a fixed schema; it never converts or infers. Questions and price notes are also LLM-written, but checked by code.
- Deterministic resolver. Every catalog item's id is built from a closed list of allowed values, for example
steel_shot|S330orstainless_sheet|304|2B|t_0_0625|w48in|l96in. Only one function can create items, so the same product can never end up with two ids, and free text never creates new items. - Thickness by curated alias, not tolerance. A stated decimal $d$ matches a stocked value $t$ only when $|d - t| \le 0.0005$ in, directly or through a curated alias. So 0.060 in becomes 16 ga (flagged as an assumption), while 0.065 in becomes a question.
- Pricing. The quote is $p = c\,(1 + m)$, where $c$ is the cheapest supplier offer that meets the buyer's requirements and $m$ is the family margin. It is compared with the benchmark $b$ as $\Delta = \frac{p - b}{b} \times 100\%$. A validator rejects any explanation containing a number that isn't among the computed fields.
- Compliance as YAML rules. Each requirement is satisfied, unconfirmed or blocking. Unknown is always shown as unconfirmed, never yes.
- Stack: Python 3.11, FastAPI, MongoDB, a React UI, Docker Compose, and BlockConvey PRISM tracing with one session per RFQ.
Challenges we ran into
- The gateway's structured output. Claude 5 models reject
temperature, the JSON-schema output mode was silently ignored, and list arguments arrived double-encoded as JSON strings. We switched to tool calling, Pydantic validation and a decoder at the boundary. - PRISM added latency. The SDK sent traces synchronously at the end of every run, so we moved delivery to a background thread that fails open.
- Wrong answers that looked right. Live runs caught three bugs our tests had missed:
- "40 pcs" of ingot plus the answer "lb" silently became 40 lb.
- "MTRs required for all alloys" blocked an abrasive line. - Alloys that don't list Mg at all were offered as matching "Mg 0.5%".
- Ambiguity is the norm. "356" could be 356.1 or A356.2; "80 grit" could be mesh or FEPA; "aluminum oxide" could be white, brown or zirconia. Each had to become a question, never a pick.
## Accomplishments that we're proud of - On a 73-line synthetic eval (emails, spreadsheets, PDF tables, government forms): 100% item match, 100% question precision and recall, 0 silent choices on ambiguous lines, and 0 invented numbers in price notes.
- Four supplier wordings ("S330 shot 50#", "Size 330", "S-330 50 LB", "SAE S330 bag") all resolve to the same item.
- A full human-in-the-loop flow that survives restarts: upload → one question message order.
What we learned
- The most valuable thing an LLM can do here is less. Restricting it to copying stable.
- "Ask, don't guess" needs structure. Numbered questions, lettered options, merged duplicates and a cap on follow-up rounds turn clarification into one tap instead of an email thread.
- Evals must target behaviour, not just accuracy. "Zero silent choices" mattered m
What's next for RUDCO-Parsing RFQ
- Real data: replace the synthetic suppliers and placeholder LME benchmark with Ru and have RudCo verify the curated size lists and compliance rules.
- Recover lost questions: near-match retrieval with LLM ranking on soft signals (buyer history, application), with the buyer still confirming.
- Harder inputs: OCR for scanned PDFs, and a review-queue UI so staff can add the solve.
Built With
- blockconvey
- langgraph
- nicrronninoo
- python
Log in or sign up for Devpost to join the conversation.