Inspiration

We started letting an agent do real errands, and it immediately needed a second agent to handle part of one. That is the moment the whole thing quietly changes: you are no longer trusting your own agent, you are trusting a stranger it picked, paid with your money, chosen from a directory where every entry looks confident and well reviewed.

People solve this by asking around. You call the one friend who used that builder. Agents have nowhere to ask, so they do what the rest of us do on a bad day: they believe the reviews. And reputation that anyone can write is worth exactly what it costs to write.

So we built the reference check. Not a rating, not a vibe score. The digging a careful person would do before handing over the money, if they had the time and the patience to open every transaction.

What it does

ou give it a job and a budget. It finds the agents offering to do the work and, for each one, asks six questions that can only be answered from public records:

  • Does its praise point at a job that was completed and paid for, by the person writing the praise?
  • Was the praise written by the same operator that runs the agent?
  • Did the payment behind the praise go back to whoever sent it?
  • How many separate operators have actually paid this agent?
  • Is the record one big customer, or several?
  • Has it handled at least as much money as you are about to hand over?

Then it either hires the one that holds up, or tells you nobody qualified and stops without spending anything.

In the demo, the agent with the most praise and the lowest price does not survive the check. It carries twelve pieces of praise; none of them stand up. Three point at payments that went back to the payer, the rest point at nothing at all. The one that gets hired carries four, each attached to a payment from a different operator that never came back. Then it does the actual job: a school newsletter goes in, and school-events.ics comes out with three dates in it, ready to open in the family calendar.

The check is careful about what it says. It reports that no completed, paid job can be pointed at, or that money returned to its sender. It never says who anyone is, whether accounts belong to the same person, or that anybody lied. Those would be guesses, and a background check made of guesses is worth nothing.

How we built it

It is a Strands agent. The six checks are registered as its tools, the agent loop is the SDK's, and the model provider defaults to Amazon Bedrock, so anyone with AWS credentials runs it unchanged with python run.py --agent.

The verdict itself is arithmetic. Counting payments, payers, and whether money came back does not need a model, and a control that stops working when a model is unreachable is not much of a control. So the same tools also run straight through: python run.py finishes on a stock Python 3.9 with nothing installed and no account of any kind.

The evidence lives on Ethereum Sepolia in the public ERC-8004 identity and reputation registries. We seeded all of it ourselves, and say so everywhere it appears: we registered the thirteen agents, sent every payment and wrote every piece of praise, including the payments that come back. It is not wild data and none of it is a finding about anyone real. What it buys is that nobody has to take our word for anything. Every transaction hash is in the repository, and scripts/from_chain.py rebuilds the whole snapshot straight from the chain so you can check the committed file against it.

Challenges we ran into

The one worth telling: our first snapshot said the wrong thing. The seeding script sent the payments that return to their sender, but never recorded that second leg, so the check could not see the money coming back and the candidate we built to fail passed everything. We only caught it because we insisted on rebuilding the snapshot from the chain rather than trusting the file the seeder wrote. The fix was to read the blocks themselves and take every transfer between those addresses, which is the only version that can see both legs of a round trip.

The second was smaller and just as instructive: the first version of the circle rule asked whether a praise author had ever been paid by an outsider, which is the wrong direction, and it quietly disqualified the honest customers along with everyone else.

Accomplishments that we're proud of

The claim that the model never decides anything has a test with teeth. We hand the pipeline a model that actively lies, one that reports everybody passed and names the wrong winner, and assert the verdict comes out identical field for field, while also asserting the liar was really called. Without that second half, the test would pass in an environment with no model at all and would have proved nothing.

The rules get the same treatment. scripts/mutation_check.py deletes each rule's condition one at a time and requires the test covering it to fail. It refuses two results that look like success and are not: a test filter that matched nothing, and a mutant that failed to import.

What we learned

Writing the rules taught us how little a reputation number carries. What survives scrutiny is not sentiment, it is money moving between parties that are not each other, in a direction that does not reverse. Everything else is decoration.

We also learned how easily a demo can flatter itself. Ours did, right up until we rebuilt the evidence from the source instead of reading our own notes.

What's next for Agent3 Vouch

The check reads one public registry today. The same six questions work anywhere the two facts exist: somebody paid, and somebody vouched. Adding a second source is a reader, not a redesign.

Built With

  • amazon-bedrock
  • erc-8004
  • ethereum
  • ethereum-sepolia
  • ics
  • python
  • strands-agents-sdk
  • web3py
Share this project:

Updates

Submission history