Inspiration

AI agents are becoming capable enough to complete meaningful professional workflows, but capability is not the same as authority. Interrupting a person for every small step destroys the value of autonomy. Giving an agent broad permission makes it difficult to prove what a person actually approved. We built VÉRTICE to preserve both: autonomous work where it is safe, and explicit human authority at the exact consequential boundary.

What it does

VÉRTICE allows a Strands agent to perform reversible preparation and then pause immediately before a consequential tool call. The Decision Inbox presents the exact proposed action, including its material parameters and a frozen action digest.

When the action is approved, only the interrupted continuation associated with that exact action may resume. The authorization is consumed once. Replaying it returns the same receipt without creating a second effect. Changing the approved amount from US$8,750 to US$8,751 is rejected. Denial produces zero effects.

The prototype deliberately labels the external business effect as SIMULATED. The Strands runtime, interrupt/resume path, browser interaction, replay check, mutation check, receipts, and evidence gates were executed and verified.

How we built it

VÉRTICE uses Python, the Strands Agents SDK, FastAPI, Uvicorn, HTML, CSS, JavaScript, Playwright, and GitHub Actions.

The canonical path is:

agent work → before_tool_call Confirm boundary → interrupt → Decision Inbox → exact resolution → same-agent resume → guarded tool → simulated effect → receipt

The browser cannot directly declare an action committed. Runtime state, the pending decision, authorization consumption, and the resulting receipt remain server-controlled.

The release process binds the evidence to one exact Git commit and source archive. Separate gates verify the real Strands application path, browser behavior, screenshots, video provenance, public-candidate safety, and final submission consistency.

Challenges

The hardest engineering challenge was proving that the interface and Strands integration were one causal execution path rather than two parallel demos. We eliminated duplicate Python module identities, unified the runtime broker, and certified the application externally through its HTTP surface.

We also had to distinguish real runtime behavior from simulated external consequences. VÉRTICE uses explicit REAL_STRANDS, SIMULATED, and REPLAY labels so the demonstration remains auditable without overstating what occurred.

Accomplishments

  • 112/112 regression tests passed in GitHub Actions.
  • 17/17 Application E2E checks passed using the official Strands runtime.
  • 15/15 browser E2E checks passed against the same application process.
  • Replay returned the same receipt with no second effect.
  • A post-approval US$8,751 mutation was blocked.
  • Denial produced zero effects.
  • Release, demo-freeze, compliance, and candidate gates all reached GREEN.
  • The final video is 3:08.7, under the five-minute limit.

What we learned

Human-in-the-loop is only the beginning. A useful approval must identify the precise action, remain tied to the runtime context that produced it, constrain what can happen after approval, and leave evidence that can be checked independently.

We also learned that honest truth labels improve the product. A judge should be able to distinguish the real agent runtime, the observed UI decision, the simulated business effect, and the replay test without reading the source code first.

What's next

The commercial path adds authenticated decision-maker identity, durable session storage, organization policy administration, provider-backed idempotency, enterprise audit exports, and production external adapters. The broader goal is authorization infrastructure for organizations adopting increasingly capable agents.

Built With

Share this project:

Updates

Submission history