Inspiration

A team ends up running a handful of AI workloads. Each one was put on a model once, by someone who has since moved on. Prices change weekly and nobody re-checks, because re-checking means pricing five workloads against ninety models and nobody has an afternoon for that.

Sumplus Steward is for the engineering leads and small-business owners who run AI workloads in production and pay that bill every month. It uses Sumplus SafeRouter's public model catalogue to review model costs against current prices.

There are two bad endings. Either nobody looks and the bill grows quietly, or a tool looks and forwards everything it finds, and now you have a second inbox.

The brief for this hackathon describes the way out of both: an agent that runs the routine work in the background and surfaces only when there is a real decision to make. Steward is built around that sentence rather than described by it.

What it does

Steward prices every workload against every live model in a public catalogue, then asks its mandate who owns each change it finds.

In the recorded demo, it applied three configuration switches with estimated savings of $4,019.60 per month, brought back one decision for a person, and left one workload unchanged because it was pinned.

Each of those five outcomes carries a sentence explaining itself, and each one leaves a receipt—including the two where the decision was to change nothing.

The mandate

The mandate is the product. It is a small file that says who owns which change:

  • Workload is pinned: Nobody touches it. Steward reports and stops.
  • Destination is not live: Refused. Production does not move onto a preview model.
  • Destination has less context than needed: Refused.
  • Saving is under $2 a month: Refused. Churn is not worth it.
  • Workload is marked quality-sensitive: Yours. Steward states the saving and waits.
  • Everything else: Steward's, and it does it.

Each workload also carries an owner-selected floor: light, mid, or flagship. Steward finds the cheapest model at or above that floor and never below it.

The bands use fixed dollar boundaries rather than quantiles of whatever is in the catalogue today, because a band that moves with the candidate list is not a promise anyone can rely on. These are pricing-based bands, not independently measured guarantees of model quality.

Today, applying a switch means changing Steward's own configuration, not migrating an external production deployment.

How we built it

Agent orchestration

The Strands Agents SDK carries the conversation through six deterministic tools for reviewing workloads, pricing them, finding cheaper options, auditing receipts, and inspecting the mandate.

The model chooses what to look at and how to say it. It never gets a vote on whether a change was inside the mandate, because that answer has to be the same whichever model is loaded—and whether or not one is loaded at all.

build_agent(model=...) takes any Strands model provider, Amazon Bedrock included.

Live pricing from Sumplus SafeRouter

Prices come live from Sumplus SafeRouter's public model catalogue, with around ninety models, current rates, and no API key required to access the catalogue.

Costs round up and savings round down, so neither figure flatters itself.

Verifiable receipts

Receipts are chained using SHA-256. Each commits to its own fields and to the hash of the receipt before it.

The verifier recomputes the whole chain from the receipts and trusts none of the stored hashes.

Web and command-line interfaces

There is a live page as well as a command line. The web page is deployed on Railway, and loading it runs the review against current prices.

The page is read-only: it applies nothing and changes no mandate.

Challenges we ran into

An agent that escalated everything

The first version escalated all five workloads, which is the failure the brief warns about: an agent that hands everything back has done nothing.

The cause was a threshold we had invented—a ceiling on how large a saving Steward could act on alone. We replaced it with a floor the owner sets per workload, which is the judgement a person should make once rather than every month.

Bands that moved with the catalogue

Our first attempt at the bands derived them from quartiles of the candidate set, and a test caught it: the boundaries moved as the catalogue moved, so the same workload could qualify one week and not the next.

Fixed dollar bands, calibrated once against the live catalogue, hold still.

A subtle receipt-verification bug

The verifier carried the stored previous hash forward instead of the recomputed one, so editing a receipt did not reach the receipt after it.

A chain where tampering does not propagate is just a list. We found it by weakening the verifier on purpose to check that the test could go red, then corrected the implementation.

Accomplishments that we're proud of

  • The review runs with no model at all. The arithmetic is arithmetic and the mandate is the mandate. A spending control that stops working when a model is unavailable is not a control, so the agent sits on top of the review rather than inside it.
  • Twelve tests cover the implementation. The chain test asserts both halves of a tampering break: the edited receipt fails its own hash, and the receipt after it fails its link. Asserting only the first would pass even if links were never checked.
  • A live review page complements the command line. Reviewers can inspect outcomes against current catalogue prices without applying configuration changes.
  • Every outcome is explained and recorded. Actions, escalations, and refusals all leave receipts.

What we learned

Refusals need sentences. "Refused" with no reason is what makes people switch these controls off, so every outcome in Steward carries an explanation. Refusals take their place in the receipt chain alongside actions.

We also learned to separate conversational flexibility from policy enforcement. The model helps users understand the review; deterministic code decides what the mandate allows.

What's next

Steward reads a catalogue and changes its own configuration today.

The same mandate shape can cover the calls themselves: a per-call spending ceiling, a host allowlist, and a receipt for every request. That is where this goes next.

Built With

  • amazon-bedrock
  • python
  • railway
  • sha-256
  • strands-agents-sdk
Share this project:

Updates

Submission history