Inspiration “We leave for our first family camping trip on Saturday. Can you get everything here by Friday evening, for under $200?” A shopping assistant can turn that request into a purchase plan. The important question comes before checkout: what happens if stock disappears, a price changes, or a payment reply never arrives?

Rehearsal is for people deciding whether to trust an agent with a transaction. It runs the proposed plan in isolated virtual environments and measures the consequences before an external order or payment can happen. The result is an impact report: what the plan spends, whether it meets the deadline, where it fails, and which alternative meets the tested conditions. Consumer shopping assistants such as Rufus inspired the example; this POC uses a synthetic catalog and has no Amazon or Rufus connection.

What it does The demo starts with one concrete request: one tent and two lanterns, delivered to the family home by Friday at 6 PM, for at most $200 including shipping. If no tent can arrive, do not buy lanterns alone.

Freeze the request and source. Capture the catalog's prices, stock and delivery schedule, together with the shopper's budget, quantities, destination and deadline. Propose a plan. Amazon Nova 2 Lite uses Strands tools to inspect that snapshot, propose a complete-basket plan and rehearse it. Measure alternatives. Execute three supplier plans across four conditions in twelve isolated worlds: as quoted, A losing tent stock, A increasing its tent price, and an injected missing payment reply. The server completes the full test set even if the agent explores only some conditions. Show the impact. An independent verifier replays each journal to calculate spending, on-time receipts, unresolved reservations and duplicate payments. The dashboard compares every result and exposes the tested coverage. Stop for a decision. The person can decline or accept a qualifying plan for execution review. Acceptance rechecks the source and report expiry, then records a downloadable brief bound to that exact plan and report. It does not place an external order or authorize a real payment. The reported numbers come from executed order, payment and delivery state transitions. They are measured simulation outcomes, not estimates generated in the model's answer.

Why it matters A low price alone does not answer the family's question. A basket that arrives on Monday misses the trip; lanterns without a tent do not satisfy the request. Rehearsal brings these consequences into the decision before money is committed.

Its intended benefit is a more informed choice about whether and how an agent should act. The current demo shows a concrete tradeoff and the evidence behind it. Reduced financial losses, time savings and improvements in user decision quality have not yet been measured.

How it was built React and TypeScript run on private S3 behind CloudFront. API Gateway and Lambda admit a bounded request, freeze the source snapshot and start a Step Functions workflow. Amazon Nova 2 Lite runs through Strands on a separate AgentCore preview runtime. Its tools can inspect, propose and simulate; its IAM role cannot invoke the commerce or seller functions or write the commerce table.

Each virtual world has its own state and hash-chained journal in temporary runtime memory. Raw evidence and the provider-usage ledger are exported to S3. A separate Lambda stops the session, independently audits all twelve cells and writes the final report outside the model-writable prefixes. DynamoDB stores source revisions, cost admission, run state and explicit decisions. A source condition check and decision write occur in the same transaction.

There is no always-on application or database server. AgentCore and Lambda execute on demand; S3, DynamoDB storage, logs and service requests remain metered. Each preview reserves $0.50 from the existing finite shared model allowance before it starts. Missing provider usage retains its reservation. Finalized report reads reuse stored evidence rather than making new model calls.

Challenges Measuring consequences: a fluent explanation is not proof of delivery. We execute a frozen plan and independently replay its journal. Missing, duplicate or corrupted cells block the handoff instead of disappearing from the denominator.

Payment uncertainty: in the response-loss simulation, payment commits before the response is hidden. The simulation retains the reservation and queries the same order, avoiding a second authorization. This is an injected virtual transport fault, not a real network-failure experiment.

Evidence going stale: a successful simulation can become irrelevant when the source changes. The report carries a source hash and a fifteen-minute validity window. A plan without qualifying evidence, a report that has expired, or a source that has changed cannot receive a new execution brief. Duplicate decisions return the original brief without extending its validity.

Honest scope: the stock and price shocks target A. B meeting all four conditions does not show that B survives its own disruptions. The report exposes these limits beside the results.

Accomplishments and results A real Google Chrome walkthrough exercised the deployed Nova planning, simulation, report, decision export and reload flow. Nova initially proposed A. The independently verified results were:

Plan Normal simulated spend Normal simulated arrival Tested conditions met A $149 Thursday, 6:02 PM 2/4 B $179 Friday, 12:02 PM 4/4 C $129 Monday, 6:02 PM 0/4 The report recommended B for review. Seven actual Nova calls used 16,929 tokens and produced a $0.006197 model-cost estimate from raw provider usage; that is not the final AWS invoice or total infrastructure cost. All twelve exported journals were re-audited locally. The workflow succeeded, the runtime was stopped, the commerce table had zero rows for this preview, and IAM policy simulation confirmed the preview role lacked commerce authority.

The local gate passed 740 Python tests and 28 browser tests. Dedicated checks cover evidence tampering, missing cells, incomplete model usage, retained reservations, idempotent decisions, expiry and a source change between validation and commit. These are implementation checks, not measured user benefit or real-world reliability.

What was learned The most useful output was the tradeoff before committing: A saves $30 compared with B but fails two tested conditions; B fits the deadline in this test set; C is cheap but late. Freezing the plan and keeping its evidence separate from the final decision makes that choice inspectable. The person can reject it or take the brief into a separately authorized execution process.

What is next The next validation is to compare simulation predictions with an independently controlled real execution adapter, then measure whether people make better decisions with the reports. A live catalog connector, production transaction integration, independent held-out scenarios and user-outcome measurements remain future work. This POC provides execution evidence, not a guarantee or a production purchase service.

Built With

Share this project:

Updates

Submission history