Inspiration
The people doing the most good often have the least time.
A food bank has volunteers to coordinate. A library has patrons to help. A school has families to support. A neighborhood organization has problems to solve. A small nonprofit may have more people depending on it than it has people available to do the work.
AI agents could give these teams another set of hands. They could review requests, sort information, handle routine cases, coordinate follow-ups, and take repetitive work off people's plates.
But there is a catch: the more useful an AI agent becomes, the more responsibility it can take on—and the more important it becomes to know what it has actually earned the right to do.
We didn't want to build another AI that simply says, "I can help."
We wanted to build a system that helps people answer:
"How much responsibility has this AI actually earned?"
That became the idea behind VOUCH.
AI agents shouldn't be trusted just because they're capable. They should earn trust before they earn authority.
We use a simple visual metaphor to make that idea understandable: a knight doesn't begin with golden armor and unlimited authority. It trains, proves what it can do, and earns greater responsibility.
In VOUCH, an AI agent starts with limited authority, proves itself through verified outcomes, and can earn more responsibility over time.
The knight earns its armor.
The armor is just the metaphor. The underlying system is real.
And the reason we're building it is bigger than the metaphor: we want community-serving teams to be able to use AI to do more good without giving AI unlimited power.
What it does
VOUCH is an AI trust and responsibility layer for groups using AI agents to help people.
VOUCH lets an AI teammate start small, prove it can handle work reliably, and earn more responsibility through verified results.
When a request exceeds that earned responsibility, VOUCH can bring in a person. When the agent fails verification, its authority can decrease. When a request violates a hard boundary, VOUCH blocks it.
The AI can help. The group decides how much responsibility it has earned.
At its core, VOUCH answers one question before an AI agent takes consequential action:
Has this agent earned the right to do this?
The end-to-end workflow is:
Request → Evidence → Agent Recommendation → Authority Decision → Action → Verification → Authority Update
VOUCH evaluates evidence, verification state, risk, reversibility, policy, context, current authority, verified performance history, and hard safety constraints.
It produces one of three outcomes:
| Decision | What it means |
|---|---|
| ALLOW | The agent has earned enough responsibility to act autonomously. |
| APPROVAL_REQUIRED | The request may be valid, but a person must decide. |
| BLOCK | The request violates a hard boundary and cannot proceed. |
The critical distinction is:
The agent can recommend. VOUCH decides what the agent is authorized to do.
Our working proof uses claims resolution as a concrete example.
The claims workflow is not the product itself. It demonstrates how the VOUCH trust model works.
The agent starts with $250 autonomous authority and successfully resolves a $124 duplicate charge.
The result is independently verified.
VOUCH increases the agent's earned authority:
$250 → $500
The increased authority isn't just a number on a dashboard. It changes what the agent is actually allowed to do.
The agent then encounters a $1,240 refund exception.
Because $1,240 exceeds its earned $500 authority, VOUCH returns:
APPROVAL_REQUIRED
A human can authorize that specific action, but the agent remains at:
$500
Human approval does not become permanent AI permission.
Later, an expected outcome fails independent verification.
VOUCH responds:
$500 → $100
The agent's autonomous capability decreases.
And even a strong performance history cannot override a hard boundary. A conflicting or untrusted instruction can be:
BLOCKED
The complete lifecycle is:
PROVE → EARN → ACT → VERIFY → ADJUST
For Good Neighbor Agents, that means AI can take on more routine work while humans remain responsible for the decisions that actually require them.
A food bank worker can spend less time processing routine cases.
A library team can automate repetitive administrative work.
A school or community organization can use AI assistance without having to choose between automating nothing and trusting everything.
VOUCH creates a middle path:
Let AI handle what it has demonstrated it can handle.
Bring people in when the responsibility exceeds what the agent has earned.
Block what should never happen.
How we built it
VOUCH is a real, end-to-end agent workflow built around a strict separation between agent intelligence and agent authority.
Strands Agents + Amazon Bedrock
The agent is built with the Strands Agents SDK and uses Amazon Bedrock for model-powered reasoning.
The Strands agent receives structured case information, uses read-only tools to inspect available evidence, reasons about the request, and produces a structured recommendation containing:
- Proposed action
- Reasoning
- Evidence references
- Confidence
- Requested authority
The most important architectural decision is:
The agent can recommend. It cannot authorize itself.
Independent authority engine
After the Strands agent produces its recommendation, the request enters a separate server-side VOUCH authority engine.
The authority engine evaluates:
- Evidence
- Risk
- Reversibility
- Policy
- Request context
- Current authority
- Verified performance
- Verification state
- Hard safety constraints
It returns:
ALLOW
APPROVAL_REQUIRED
or
BLOCK
The model's recommendation is therefore an input to authorization—not authorization itself.
Increasing model capability does not automatically grant increasing permission.
Enforced safety boundaries
The agent cannot use its own reasoning to increase its authority.
It cannot:
- Grant itself more permission
- Authorize its own consequential action
- Modify its own trust state
- Modify audit state
- Convert a human approval into permanent autonomous authority
These boundaries are enforced through the application architecture and server-side state rather than relying on a prompt telling the model to behave.
The knight cannot put on its own armor.
Human authorization
When VOUCH returns APPROVAL_REQUIRED, the system creates an explicit human authorization path.
The approval is associated with the specific action, case, evidence, and session.
Human-authorized execution is recorded separately from autonomous execution.
Most importantly, approving one exceptional action does not increase the agent's autonomous authority.
Human decision ≠ permanent AI permission.
Independent verification
VOUCH does not assume that a successful API call means a successful outcome.
After execution, the system independently verifies the expected result against durable state.
That verification becomes evidence for the next authority decision.
The feedback loop is:
Recommend → Authorize → Execute → Verify → Adjust → Reevaluate
An agent earns more responsibility through verified outcomes—not simply through activity.
Durable state with Postgres
VOUCH persists case, action, authorization, trust/authority, verification, and audit state using Postgres.
This matters because authority is behavioral, not cosmetic.
When verified performance changes authority from:
$250 → $500
or:
$500 → $100
that persisted state changes what the agent is allowed to do on subsequent requests.
The state isn't decoration. The state controls the agent.
Honest AWS failure handling
When Bedrock is unavailable, VOUCH does not pretend a model response came from AWS.
The application falls back to a clearly labeled deterministic evaluator so the complete workflow remains demonstrable without misrepresenting the source of the decision.
The fallback is visible. The AWS-backed result is never fabricated.
Design
VOUCH's interface is designed to make a complicated trust model understandable at a glance.
Users can see:
What the agent knows.
What the agent recommends.
What the agent is allowed to do.
Why VOUCH made that decision.
What happened afterward.
How that outcome changes future authority.
The knight progression makes the concept intuitive: an agent starts with limited responsibility and earns its armor as verified performance demonstrates that it can responsibly take on more.
But the visual progression is not a gamified score disconnected from permissions.
The underlying authority actually changes.
That makes VOUCH approachable without hiding the seriousness of the decisions being made.
Challenges we ran into
The hardest part wasn't getting an AI agent to produce a recommendation.
It was answering:
What happens after the agent makes the recommendation?
We discovered that these are fundamentally different questions:
What does the AI think should happen?
What is the AI actually allowed to make happen?
A prompt cannot reliably serve as an authorization boundary.
A confidence score cannot serve as permission.
And a human approval button should not accidentally become permanent AI authority.
That led us to the core architecture:
Agent intelligence on one side. Independent authority on the other.
We then had to make "earned autonomy" real.
If a number changes but the agent's actual capabilities don't, nothing has been earned.
So VOUCH ties authority directly to future behavior:
$250 → $500 after verified success.
$500 → $100 after verification failure.
The change in authority actually changes the outcome of the next request.
We also had to separate human responsibility from agent capability.
A person can authorize an exceptional $1,240 action without teaching the agent that it has earned $1,240 of autonomous authority.
Finally, we needed to make an abstract trust model understandable.
Instead of simply explaining it, we made the system demonstrate it:
$250 → $500 → $100
ALLOW → APPROVAL_REQUIRED → BLOCK
The judge can see the system change what the agent is actually allowed to do.
That became one of our biggest design lessons:
The state isn't decoration. The state controls the agent.
Accomplishments that we're proud of
We made human control enforceable.
VOUCH doesn't just say "humans are in the loop."
It defines when humans need to be involved, what the agent can do independently, what evidence supports additional responsibility, what happens after failure, and what can never be overridden.
We built a non-trivial Strands agent workflow.
Strands Agents and Amazon Bedrock power genuine agent reasoning inside a complete workflow that includes tools, structured recommendations, independent authorization, action execution, human approval, verification, durable state, and adaptive authority.
The model is not the authorization system.
That separation is the technical idea behind VOUCH.
Earned responsibility changes actual capability.
The proof is:
$250 → $500 → $100
The agent's authority isn't a visual score.
It changes the outcome of the next request.
We made verification part of trust.
An agent doing something successfully is not enough.
VOUCH asks whether the intended result actually happened.
Verified outcomes become evidence for future authority decisions.
We built for people, not just for AI.
The Good Neighbor Agents track is about groups of people.
VOUCH is designed for teams such as nonprofits, food banks, schools, libraries, neighborhood organizations, and other community-serving groups.
The goal isn't to replace the people doing the work.
It's to give them another set of hands without giving those hands unlimited power.
We made responsible autonomy understandable.
The knight metaphor gives people a simple mental model:
Train. Prove. Earn responsibility.
The cute armor progression helps make the concept memorable, but the important part is what happens underneath the metaphor.
The agent's authority is enforced by the system and can increase or decrease based on verified performance.
We built around a real human benefit.
For a small community-serving organization, the value of AI isn't another generated answer.
It's time returned to people doing the work.
Our seeded workflow estimates approximately 14 minutes returned for an autonomous verified resolution, compared with approximately 6 minutes for a human-authorized resolution.
The larger vision is:
More routine work handled by an AI teammate. More human attention available for the people who need it.
What we learned
We started VOUCH thinking about trust.
We ended up thinking much more about responsibility.
An AI agent does not have to be perfectly trustworthy before it can be useful.
It needs boundaries that match what it has actually demonstrated.
That led us to a different view of autonomy:
Autonomy is not a switch. It is a spectrum of responsibility.
An agent can start small, demonstrate that it can handle a type of work, earn more authority, lose that authority, and encounter situations where a person must decide.
We also learned that verification matters more than activity.
An API call succeeding does not prove the intended outcome happened.
And human-in-the-loop does not have to mean human approval of everything.
If people must approve every routine action, the agent hasn't meaningfully increased the team's capacity.
The better question is:
Which decisions genuinely need a person?
VOUCH is our attempt to answer that dynamically.
We also learned that responsible AI doesn't have to mean incapable AI.
The goal isn't to keep an agent permanently constrained.
The goal is to give it a path to earn more responsibility without allowing it to decide its own limits.
Capability can grow.
Responsibility must be earned.
The knight metaphor helped us understand this too.
A knight doesn't receive armor because someone says it is trustworthy.
It earns that armor by demonstrating that it is ready for greater responsibility.
For VOUCH, that translates into a system where:
Trust comes from evidence.
Authority comes from verified performance.
Failure can reduce authority.
Human approval handles exceptions.
Hard boundaries remain hard.
That is the kind of AI teammate we want to build for people doing important work.
What's next for VOUCH
Today, claims resolution is our proof point.
Next, we want to apply the same trust-and-responsibility model to AI teammates helping the people who keep communities running.
Potential applications include nonprofit operations, food-bank logistics, school and family coordination, library services, community requests, and local organization administration.
The opportunity is bigger than automating one workflow.
What if a small community team could have an AI teammate that became more useful as it proved itself—without ever becoming the one that decides its own limits?
That is the future we want to build.
A system where AI can take on more of the repetitive work.
Where trust is earned through evidence.
Where capability can increase when performance proves it deserves to.
Where capability can decrease when it doesn't.
Where exceptional decisions come back to people.
And where hard boundaries remain hard.
VOUCH isn't trying to replace the people doing the work.
It's trying to give them another set of hands they can trust.
Because the goal of responsible AI isn't to keep AI from helping.
It's to make sure that as AI becomes more capable, people remain in control of how that capability is used.
Let it earn its armor.
Let the AI earn its place on the team.
PROVE. EARN. ACT. VERIFY. ADJUST.
VOUCH — AI that looks out for people.
Built With
- agentic-ai
- ai-agents
- amazon-bedrock
- amazon-web-services
- generative-ai
- postgresql
- strands-agents-sdk
Log in or sign up for Devpost to join the conversation.