What it does

Some decisions come back every single day, and the data that should settle them is scattered enough that people end up going on instinct. Not because they are careless, but because checking properly costs more time than the decision is worth, every morning, forever.

This agent takes one of those decisions and does the checking.

The first case we built it for: a Korean produce farmer weighing where their crop has been fetching more. You ask "peaches, 28th — how did the markets compare?" and it ranks them by price per kilogram on that day's auction records, with how many records each figure rests on.

You ask in Korean and it answers in Korean. The sample output in the repository is untranslated for that reason: it is the real thing, not a mock-up. The data is a Korean government feed and the user is a Korean farmer; an English demo would have been easier to read and would have been a different product.

To be exact about what that is: it compares observed prices, not shipping outcomes. It does not know your variety, your grade, your freight cost, the commission at each market, or whether you can still get a lot in tomorrow. Those decide the actual choice. This narrows the field and shows the evidence; the shipper still decides.

When the evidence is not good enough, it does not answer. It says which markets fell short, what fraction of the day it actually saw, and what would make the question answerable. That refusal is the feature. A daily decision does not need a confident guess — it needs to know when the data is thin, because tomorrow it will have to decide again.

Who it's for

People who make the same judgement call every morning with data they cannot fully check. We built for one of them first, and picked the one we could actually see from the inside.

In Korea, produce moves through 33 public wholesale markets, and for the same crop on the same day the gap between the best and worst market ranged from 45% to 5,500% across the 16 date × product pairs we measured (median 207%). Those pairs are the ones where at least two markets had three or more records each; thin markets are excluded, not averaged in. Package size is only one driver of that spread; variety, grade and origin move it too, and we did not separate them.

But that number needs a caveat we built into the product: the widest gaps are not opportunities. A 1 kg retail bag in one market and a 12 kg bulk box in another are both "cabbage," both real auctions, and comparing them per kilogram produces a nonsense spread. The agent detects that and says so, instead of pointing you at a market you cannot actually sell into.

The auction data is public. Reading it correctly is the hard part.

This is not a hypothetical user. One of us runs operations at a wholesale market corporation, and this decision is made in that building every day. We did not go looking for a problem to solve; we already had one, and we had spent years watching people solve it by feel.

The shape generalizes: a contractor pricing a job, a teacher deciding who needs help this week, a shop owner choosing what to reorder. We are not claiming those. We built the one we know.

How it works

Built on the AWS Strands Agents SDK. The model reads the question, picks a tool and fills in its arguments — turning "peaches, 28th" into a product and a settlement date. (The feed is keyed by settlement date, which is not always the day a lot was shipped. We query and report that date, and we do not convert between the two.) Everything after that is code: the verdict is computed deterministically, and the sentence the user reads is written by the tool, not by the model. If the right tool never ran, the agent says so instead of answering.

That split is the whole design. Here is why we made it.

We audited a tool we had already published — and it was answering on 0.77% of the data.

What we counted On 2026-08-28
Auction records in one day 129,536
What the tool fetched before filtering by product 1,000
Distinct markets in that slice 6 of 33 — a single market was 64.3% of it
Markets absent from that slice 27 of 33
Of those, ones we confirmed it never reaches 26 of the 32 that actually traded — every market except the six above. Among them Seoul Garak, 25.6% of the day on its own (33,124 of 129,536), 2.6x the next market

All figures above: data.go.kr public auction feed (Korea Agro-Fisheries & Food Trade Corp.), settlement date 2026-08-28. The 33 is that feed's roster; 32 is how many actually traded produce that day. The one difference is a seafood market. Counts as measured 2026-08-30; the feed is revised retroactively, and the same script returned 129,704 for the same date on 09-01, with the ratios unchanged. We give the number and the day we measured it, because a reviewer re-running it will get a slightly different total.

It still produced a confident "nationwide market comparison." We have audited exactly one such tool — our own — so we make no claim about anyone else's. What we can say is what it cost us: the data looked thin, and the data was not thin. The tool was. We spent a day believing the wrong thing.

So the agent measures its own evidence first:

  • Coverage: what share of the day did we actually retrieve? A truncated sample cannot support "market X pays the most," because a market we never saw might pay more and we cannot rule it out.
  • Unit normalization: auction prices are quoted per package, and packages are 3, 4, 5, 10 kg. Averaging raw prices compares a small box to a big one. Correcting this changed which market ranked first in 9 of 16 date × product pairs we tested.
  • Outliers: one record at ₩500,005/kg among 656 moved that market's average by 21%. We report the median and say so when the mean disagrees.
  • Package class: normalizing to ₩/kg still does not make two markets comparable. A 1 kg retail bag and a 12 kg bulk box are both "cabbage," both real auctions, and their per-kilogram gap reads as 5,500%. The agent shows each market's typical pack size and says when the top and bottom are different kinds of trade.
  • Sufficiency: minimum markets, minimum records per market. Fail any of these and the tool returns insufficient, with reasons.

Asking a model to be careful is a hope. A tool that returns insufficient is a constraint, and a constraint is not a guarantee. On the normal path the design leaves the model nothing to answer with; our own measurements below show what still leaks past it, and how much.

Three things we had to learn the hard way, each measured before and after:

  • A refusal screen must not show the numbers. We first listed prices under "reference only, do not cite", and the model cited them in 4 of 6 runs. A label is a request. Removing them is a constraint. 1 of 6 after.
  • The line the model copies must carry no statistics. Our refusal headline once quoted a coverage ratio; the model spun those digits into a table of per-date record counts it had never queried. Plain words instead: 0 of 6.
  • The answer sentence is written by code. Letting the model rephrase it put Vietnamese and Chinese characters into Korean output in 4 of 6 runs, and once relabelled a quantity as a record count. The model now chooses the tool; the tool writes the sentence; if the right tool never ran, the agent refuses. Across the three demo scenes, 4 runs each, 12 of 12 produced the expected verdict line with no non-Korean script in it. That is a small sample on three fixed questions, not a reliability figure.

A tool that crashes is the same problem in a different coat. When ours hit a day with no data, it raised, and the model filled the silence with markets that do not exist. The tool now converts any exception it reaches into a refusal, and our tests enforce that. We have verified the paths we found; we are not claiming we found every one.

A web UI, on the same path. web.py serves a one-page interface that calls the same agent.run() the command line uses; there is no separate UI path. The table on screen is built from the Evidence object the tool computed on that very call, so the numbers a person sees and the sentence the model was handed come from one computation. On a refusal the page shows the reasons and the markets it saw and withholds the prices, for the same reason the text output does. The first line under the input states which dates and products the bundled sample can answer, so a demo cannot be mistaken for coverage it does not have. Screenshots of an answer and a refusal, both from real model runs, are in the repository README.

What we're honest about

  • The flip-rate result is 16 pairs across 2 full days. We say "often flips", not a percentage.
  • The market-bias measurement is one day. We did not test whether the ordering is stable.
  • The sufficiency thresholds are chosen, not derived.
  • The before/after model figures are 6 runs each, one local model, one machine.
  • Two claims we made about our own system did not survive being measured, and we withdrew them — a timing figure that turned out to be one draw from an 18–46s spread, and a paraphrase problem our own fix had already solved. Both retractions are in the repository history.
  • Every number above is reproducible from measurements/, which maps each script to the claim it supports and the sample size behind it.

We think that section belongs in the submission, not in a footnote. An agent that reports its own limits should come from a team that reports theirs.

New work, and what pre-existed

Every file that ships was written during the hackathon window: evidence.py, data.py, tools.py, agent.py, config.py, both test suites, and everything in measurements/.

One thing pre-existed and we are disclosing it because the story above depends on it: korean-agriculture-mcp, an MIT-licensed MCP server for the same public data, published by the same author before this hackathon. No code from it is reused here — it is the tool we audited, the one that was answering on 0.77% of a day. It is named so that the 0.77% figure has a subject you can go and check.

The demo data file is an extract of the public data.go.kr auction feed, built by the script in this repository. It contains six public fields (product, variety, quantity, price, unit weight, market name) and nothing else.

Bonus blog posts

Three posts on builder.aws.com, all published before the deadline:

Built With

  • data.go.kr
  • mcp
  • ollama
  • python
  • strands-agents
Share this project:

Updates

Submission history