Inspiration

One kitchen sells on three Korean delivery platforms — Baemin, Coupang Eats, Yogiyo. Every month three settlement statements arrive, each with its own format and its own commission rate. Reconciling them line by line takes hours the owner does not have, so a few hundred won of over-deduction per order is simply never contested.

The obvious fix — "have an LLM read the statement and tell me what's wrong" — is worse than doing nothing. A model that hallucinates one commission figure, or files a dispute on a hunch, costs the owner a relationship with a platform they depend on for revenue.

So the design constraint came first: an agent that is allowed to be useful only where it can be certain, and is required to stop everywhere else.

What it does

Settlement Sentinel runs in the background and reconciles a month of settlement data per platform:

  1. Recomputes every fee from the owner's contract rate — never the rate printed on the statement.
  2. Classifies each gap by evidence type, and assigns confidence from a fixed Python table.
  3. Acts or refuses. At or above 0.85 confidence it files the dispute itself. Below it, it files nothing and asks the owner one question.
  4. Records the owner's answer and never re-asks. Re-running a platform returns skipped_already_filed rather than filing twice.

The demo has three outcomes across three platforms:

Platform Finding Confidence Action
Baemin BM-1005, internally inconsistent line 0.90 filed automatically
Coupang Eats CE-2004, billed 10.9% against a 9.9% contract 0.95 filed automatically
Yogiyo YG-3005, 3,000 KRW gap with no stated cause 0.55 refuses — asks the owner

Note the third row. YG-3005 is the largest amount in the dataset and the agent is least confident about it, so it stops. Confidence and money move in opposite directions, and the gate only looks at confidence. That is the whole design.

How we built it

Built with the AWS Strands Agents SDK on Amazon Bedrock (Claude Sonnet 4.6, us-west-2).

Three deterministic Python files run in a fixed order, each refusing to run unless the previous one has:

  • tools_settlement.py — parse, recompute, flag
  • gate.py — decide the action. The only place an action is authorized
  • report.py — render the report and the owner queue

agent.py holds the Strands agent. It sequences the tool calls and writes a two-sentence takeaway. It does not compute a figure, does not decide an action, and cannot skip the gate.

Every number is computed, gated, and rendered by deterministic Python. The LLM only orchestrates.

Challenges we ran into

Circular verification. My first version recomputed each commission using the rate printed on the statement being audited. Against a statement billed at 10.9% on a 9.9% contract, every row reconciled to zero and the overcharge passed as clean. The arithmetic was correct; the audit was worthless. A check that consumes the claim it is testing cannot fail, and a check that cannot fail is not a check. The fix is a contract rate table held on the owner's side — an independent second source. (I wrote this up in detail here: https://builder.aws.com/content/3IiguE9Tk9MsRWolGzNyjD70ZA7/agents-for-humans-my-audit-agent-passed-a-109percent-overcharge-as-clean)

A claim that did not match the code. I originally wrote "the LLM never touches a number." A cross-check showed that was false: tool results including figures do enter the model's context, and one code path was literally instructing the model to transcribe amounts. I moved that rendering into Python and rewrote the claim to something the code actually supports. Overclaiming in a verification project is self-defeating.

Judges without Bedrock access. A reviewer whose account cannot reach the model would see nothing at all. So --no-llm runs the same three files in the same order with no model in the loop — no credentials, no network. It doubles as the honest test of the central claim.

Accomplishments that we're proud of

The claim on the README is falsifiable by running the code:

python agent.py baemin           > with_llm.txt
python agent.py baemin --no-llm  > without_llm.txt
diff with_llm.txt without_llm.txt

Every table row is identical; only the tool trace and the model's closing sentences differ. If a row ever differed, the claim would be false and should be treated as such. That line is in the README.

I also ran ablations rather than assuming components mattered. Remove the contract rate table and every discrepancy vanishes. Skip the gate and the model narrates a problem, invents a confidence score of "70%" that exists nowhere in the code, and files nothing. Both results are why those parts are there.

What we learned

  • Ask where your comparison value came from. If it traces back to the artifact under test, the check is a ritual.
  • Confidence should be a property of the evidence, not of the amount. Tying them together is how an agent talks itself into acting on the expensive uncertain case.
  • Prompts are not enforcement. Told not to transcribe figures, the model sometimes did anyway. Anything that must hold is in Python.
  • State limitations first. The README lists six, including that filing is simulated and the data is synthetic. A reviewer finds them anyway; being first is cheaper.

What's next for Settlement Sentinel

Real platform partner APIs instead of a simulated ledger. Composite evidence, so a rate mismatch and an unrelated settlement gap on the same order are scored together rather than by the dominant type. Cross-month detection, where the same order id reappears with a different adjustment. And a deployment on Amazon Bedrock AgentCore so the reconciliation runs on schedule rather than on command.

Built With

Share this project:

Updates

Submission history