Inspiration

AI agents are getting good enough to do real business work: reconciling invoices, researching exceptions, updating records, and coordinating across systems.

What kept bothering us was what happens after the reasoning is correct.

A finance agent might correctly understand a vendor request. A human might legitimately approve a bank-account change. But the action that eventually reaches the downstream system can still differ from what the person approved.

That creates a gap that ordinary human-in-the-loop systems do not fully solve.

For this project, we wanted to explore a stricter model:

A human approval should authorize one exact consequential change, not create a general permission to continue.

We chose a vendor and payment exception workflow because it makes that problem concrete. An email can request that a vendor's banking information be changed, but the email itself should not be enough to make that change authoritative. And if a human approves a change to one destination, that approval should not silently authorize a different destination later.

That became the center of the project:

Correct reasoning + valid human approval still does not authorize the wrong effect.


What it does

Admissible Finance Exception Operator handles a synthetic vendor-invoice case from ordinary work through a consequential downstream change.

The agent first does the work that does not require a person. It matches the invoice to the purchase order, checks the freight terms, identifies the requested bank change, and prepares the case.

If evidence is missing, that remains agent work. The agent—or another agent assigned to that task—can gather or verify the missing information. A person is not interrupted simply because the case is incomplete.

Only once the evidence is sufficient does the workflow reach the actual human boundary.

The user is then asked one exact question:

Approve the vendor-bank change from 4812 to 9371?

If the user approves it, the vendor record still does not change immediately. Instead, the decision creates short-lived, single-use authority for that exact change.

The demo then deliberately changes the downstream request to 4812 → 9351.

That request is blocked.

The original 4812 → 9371 change can still execute, and it can execute only once. A replay produces no new effect.

All data in the demo is synthetic, and the project does not release payments.


How we built it

The professional agent is built with the Strands Agents SDK and uses Amazon Nova Lite for model reasoning.

Amazon Bedrock AgentCore Runtime hosts the managed agent path, and AgentCore Memory carries the case forward across sessions.

We deliberately separate memory from trust. Persisting an observation does not automatically make it authoritative business state. The agent can preserve context while the workflow still decides what information is allowed to participate in consequential decisions.

Our shorthand for that became:

Persistence is not admission.

When enough evidence exists, the agent reaches a native human-decision boundary. Approving the exact change creates a single-use authority tied to the current case, vendor, effect type, old value, new value, and execution context.

The actual downstream effect is enforced separately. AgentCore Gateway + Policy provides an AWS enforcement layer, while Admissible keeps the executed effect bound to the human decision that authorized it.

For the hosted judge experience, we built a secure browser flow with a shared demo access code, isolated browser sessions, fresh disposable demo runs, and seven explicit AWS_IAM-protected API Gateway routes. The browser never receives AWS credentials or authority material.

The final demo follows one continuous case rather than a collection of disconnected scenarios. The same case moves through ordinary work, evidence gathering, human decision, blocked mismatch, exact execution, replay protection, and audit proof.


A key design decision: evidence gaps vs. authority gaps

One of the most useful ideas that emerged while building the project was distinguishing two very different reasons a workflow might stop.

If evidence is missing, that should usually create more agent work.

If human authority is missing, that should create human work.

In practice:

ordinary work        → agent
missing evidence     → agent gathers or verifies
human authority      → ask a person
exact execution      → enforce what was approved

That distinction helped us avoid turning every uncertain case into a human escalation.

The human is brought in only when the remaining decision is genuinely theirs to make.


Challenges we faced

Keeping approval bound to the eventual effect

The hardest conceptual problem was making sure approval did not become a generic "continue" signal.

We wanted the system to preserve the exact relationship between:

  • the evidence,
  • the proposed change,
  • the human decision,
  • the authority created by that decision,
  • and the effect that eventually executes.

That is why the demo intentionally approves 9371 and then attempts 9351. The changed request is where the difference between ordinary HITL and exact effect control becomes visible.

Managed memory is eventually visible, not instantly authoritative

During final qualification, we found an intermittent issue where a Memory record had been written successfully and could be retrieved directly, but it had not yet become visible through the indexed confirmation path within our original confirmation window.

The system failed closed: the vendor stayed at 4812, no human decision was recorded, no authority was created, and no effect executed.

We hardened the confirmation path with a bounded read-only retry window while preserving the original safety rules: one write, exact record identity, exact content and receipt validation, and fail-closed behavior if confirmation never arrives.

AWS authorization on the judge path

We also spent a surprising amount of time on the hosted judge transport.

Our first design used an IAM-protected Lambda Function URL between the judge bridge and the application boundary. Even after extensive policy simulation, resource-policy inspection, identity verification, and a narrowly scoped CloudTrail experiment, the live execution-role request continued to receive an AWS 403.

Rather than weakening the IAM boundary, we replaced that transport with an AWS_IAM-protected API Gateway HTTP API using seven explicit routes.

That path qualified cleanly and ended up being a better fit for the application.

Making the demo understandable

Another challenge was presentation.

The underlying system has several layers of evidence, authority, memory, policy, and effect enforcement. We did not want the judge experience to become a governance dashboard.

The final UI stays on one case page. Each deliberate action advances the real workflow, while the current state remains visible long enough to inspect. The audit details are there for judges who want to go deeper, but the main story stays simple.


What we learned

The biggest lesson was that human approval is not the same thing as human authority.

An approval only becomes meaningful if the eventual action remains bound to what the person actually decided.

We also learned that long-running agents need a clearer separation between memory and authority. Agents should be able to preserve observations, continue work, and gather missing evidence without silently turning every persisted claim into trusted business state.

And we came away with a much stronger view of what "Agents for Humans" can mean.

The goal is not to put a human into every uncertain loop.

The goal is to let agents do as much useful work as they safely can, then ask a person only when the missing ingredient is something only a person should provide.


Pre-existing work and hackathon work

Admissible is a pre-existing proprietary runtime/service developed before this hackathon. Its underlying implementation is not included in the public repository.

The pre-existing runtime provides capabilities for authority evaluation, evidence handling, exact delegation, effect enforcement, and reconciliation.

For this hackathon, we built the finance exception application and AWS-native experience around that runtime, including the Strands/Nova agent workflow, AgentCore Runtime and Memory integration, selective human interruption, vendor/payment case flow, AWS enforcement integration, hosted interactive judge experience, secure browser sessions, API Gateway IAM transport, synthetic downstream execution, replay protection, and browser-visible audit proof.

The public repository contains the hackathon application layer that is safe to publish, while the complete system can be evaluated through the hosted demo.


Where we would take it next

The finance workflow is intentionally narrow, but the underlying pattern is broader.

Any system where an agent can gather evidence, propose a consequential change, obtain human authority, and eventually affect an external system has the same question:

Did the thing that became real still match what was actually authorized?

Vendor-master changes are one example. The same pattern can apply to payments, access changes, infrastructure operations, procurement, healthcare workflows, or other agentic systems where approval alone is not enough.

For this hackathon, we focused on proving that pattern deeply in one case rather than building a broad collection of demos.

Built With

Share this project:

Updates

Submission history