Inspiration

Shops here buy from the same four or five distributors, and every few months a new price list lands. A CSV exported from Excel, a table pasted out of a PDF, sometimes a photograph of a printout. Two hundred lines, headers in whatever language the supplier uses, prices written 2,28 or 2.28 depending on who typed them.

Nobody compares it against the last one. There is one person doing the ordering and they have a shop to run. So a supplier moves forty items up six percent and it goes unnoticed for a year. The margin bleeds and nobody can say when it started.

The comparison itself is not hard. It is tedious, repetitive, and has to happen every single time a list arrives, which is exactly the shape of work worth handing to an agent.

What it does

Give Ratchet the list that just arrived and it:

  1. Reads it - whatever the delimiter, whatever the header language, either decimal convention. Rows whose price cannot be read are reported as skipped, never guessed at.
  2. Diffs it against the price book, treating movement under 0.5% as rounding noise rather than news.
  3. Prices the damage - ranks every movement by what it costs per year at real sales volume.
  4. Checks margins - flags the lines now below target, and separates the ones this list broke from the ones that were already under.
  5. Drafts the reply - item code, old price, new price, percentage, volume, and one specific question.

It does not overwrite the price book unless you approve it. Committing silently would destroy the very history that makes the next comparison possible.

On the bundled sample, one supplier and fourteen lines, it finds 4,417 EUR a year of silent cost increase and seven lines newly below their margin target.

The ranking is the whole point. The milk moved 4.2%, the smallest rise on the list, and it is the second most expensive thing that happened, because the shop sells 25,200 units of it a year. Sorting by percentage puts the wrong line at the top of the page. Sorting by money does not.

How I built it

Strands Agents SDK, six tools, and a hard rule: every number comes from code, and the model never does arithmetic.

The model chooses which tool to call and writes the message to the supplier. That is all it does. Each margin, percentage and annual figure is produced by core/, which runs on the Python standard library with no model, no network and no key.

This is not a cost saving, it is a usability requirement. A buyer challenging a supplier has to point at the rule that fired. "The model calculated 8.6%" does not survive that meeting. "Your list moved 2.10 to 2.28, we bought 4,080 units last year, that is 734.40 EUR" does.

The six tools:

Tool Returns
read_price_list Parsed rows, plus how many were skipped as unreadable
compare_to_price_book Every SKU that rose, fell, appeared or was withdrawn
price_the_damage Findings ranked by annual impact, with totals
check_margin What one cost does to one product's margin, and the shelf price that fixes it
buying_history Current cost, shelf price, volume and target for one SKU
commit_price_book Accepts the new list as baseline, only after approval

Strands is provider-agnostic, so Ratchet is too. RATCHET_PROVIDER selects Bedrock, Ollama, OpenAI or Anthropic and nothing else changes. Ollama matters here: the people this is built for are not going to open a cloud console, and it runs free and local with no account at all.

There are two implementations of the decision logic. core/ in Python is canonical - it is what the agent calls. web/core.js mirrors it so the live page runs the real pipeline in your browser with nothing installed. tests/parity.mjs runs both over the same inputs and asserts the outputs are identical, so the demo page and the agent can never quote different figures.

Challenges I ran into

"Margin break" was telling the buyer the wrong thing. The first run flagged lines as broken when they had already been below target before the new list arrived. That is a pricing problem the shop gave itself, not something the supplier just did, and taking the wrong one of those into a supplier meeting wastes the meeting. Findings now carry already_below and newly_broken separately, and the wording differs.

My own sample data was flattering the product. I had set shelf prices so low that nine of ten lines sat under target no matter what the supplier did, which made the demo look better than the tool was. Rebuilt the sample shelf prices from real target margins so the increases visibly cause the breaks.

Real price lists are not clean. Semicolons, tabs, commas. Headers in Albanian or English. "€ 1.234,56" and "1,234.56" meaning the same number. The parser decides the decimal separator by which of . or , appears last, which holds for both conventions, and it drops rows it cannot read rather than inventing a price.

The build environment had no package index. I could not install the SDK where I was writing the code, so the decision layer was written dependency-free from the start and covered by 38 assertions that run on the standard library alone. That constraint produced a better architecture than I would have chosen freely.

Accomplishments that I'm proud of

Two implementations of the same logic, in different languages, held to byte-identical output by an automated check. It means the live demo is not a mock-up of the agent - it is the agent's arithmetic, running in your browser.

And the honesty of the output. Every flagged line shows the sentence-level reason it was flagged, the volume behind it, and the shelf price that would fix it. Nothing asks to be trusted.

What I learned

That the interesting part of an agent is usually not the model. The tool boundaries, the extraction contract, the decision to keep arithmetic out of the language model - that is engineering, and the system is better for the model being replaceable.

What's next for Ratchet

Reading price lists straight out of email, a PDF pipeline so lists nobody has pre-structured still work, scheduled runs with a weekly digest, and multi-supplier comparison - the same item across every distributor who sells it.

Built With

Share this project:

Updates

Submission history