AgentVerify
Agents can hire agents. Who verifies the work?
Inspiration
The rise of autonomous agents is creating a new agent economy where AI agents can discover services, communicate with each other, delegate tasks, and make decisions.
However, one important problem remains:
How does an agent know whether another agent actually delivered what was requested?
A provider agent can return a response that appears correct but may:
- Miss important requirements
- Provide incomplete information
- Return invalid structured data
- Make unsupported claims
- Contain contradictory information
- Refuse part of a task without clearly communicating the limitation
Traditional systems generally focus on whether a request was executed, not whether the result actually satisfied the original agreement.
This inspired us to build AgentVerify, a verification and reputation layer for autonomous agent-to-agent transactions.
Our core idea is simple:
Do not trust the response. Verify the delivery.
What it does
AgentVerify is a callable A2A service that independently evaluates a provider agent's delivery against the original request, agreed requirements, expected output structure, and supporting evidence.
The verification pipeline consists of multiple checks.
1. Requirement Verification
AgentVerify breaks the agreement into individual requirements and checks each requirement separately.
Instead of simply asking whether the response "looks correct", it determines:
Requirement 1 → Satisfied
Requirement 2 → Satisfied
Requirement 3 → Missing
Requirement 4 → Partially satisfied
This makes the final decision explainable.
2. Completeness Verification
The system identifies missing:
- Required fields
- Deliverables
- Information
- Evidence
- Output components
This allows AgentVerify to distinguish between a complete delivery and a partially completed task.
3. Schema Verification
When a structured output is expected, AgentVerify validates the provider response against the required JSON schema.
This verifies whether:
- Required fields exist
- Data types are correct
- Structure is valid
- Constraints are satisfied
4. Evidence Verification
For tasks requiring evidence, AgentVerify checks whether important claims are supported by the provided evidence.
Unsupported or weakly supported claims are flagged instead of automatically being accepted.
5. Consistency Verification
AgentVerify checks for contradictions within the delivered information.
For example:
Claim A: Product availability = 25 units
Claim B: Product availability = 0 units
The system identifies the conflict and lowers the verification confidence.
6. Usability Verification
A technically valid response is not necessarily useful.
AgentVerify evaluates whether the delivered result contains the information required for the requesting agent to actually use it.
7. Safety and Refusal Verification
The system recognizes legitimate refusals and insufficient information.
It does not treat every refusal as a failure.
It can distinguish between:
Successful delivery
Partial delivery
Failed delivery
Valid refusal
Insufficient information
Verification Score
AgentVerify uses a transparent weighted scoring model.
$$ S = 0.30R + 0.20C + 0.15T + 0.15E + 0.10K + 0.05U + 0.05A $$
Where:
- (R) = Requirement Coverage
- (C) = Completeness
- (T) = Structure
- (E) = Evidence
- (K) = Consistency
- (U) = Usability
- (A) = Safety
The score is intentionally transparent so that another agent can understand why a provider received a particular result.
Score vs Confidence
AgentVerify separates verification score from verification confidence.
The score represents how well the delivery satisfies the requirements.
Confidence represents how reliable the verification decision is based on the available information and evidence.
For example:
Score: 92/100
Confidence: 96%
Score: 78/100
Confidence: 54%
A high score with low confidence indicates that the available information may not be sufficient to fully trust the result.
When evidence is insufficient, AgentVerify can return an uncertainty state instead of inventing a conclusion.
Verification Verdicts
AgentVerify produces clear machine-readable outcomes.
| Verdict | Meaning |
|---|---|
VERIFIED |
Delivery satisfies the agreed requirements |
PARTIAL |
Important parts were delivered but some requirements are missing |
FAILED |
Delivery does not satisfy the requirements |
REFUSED |
Provider refused or could not safely perform the task |
UNKNOWN |
Available information is insufficient for a reliable decision |
It also provides a recommended action:
ACCEPT
ACCEPT_WITH_CAUTION
REQUEST_REVISION
REJECT
ESCALATE
INSUFFICIENT_DATA
How we built it
AgentVerify is built as a full-stack callable service.
Technology Stack
- Next.js
- TypeScript
- Tailwind CSS
- Node.js
- PostgreSQL
- Prisma
- JSON Schema
- AJV
- SharedOS
- SharedNet
- REST APIs
- SHA-256
- Vitest
Architecture
The main workflow is:
$$ \text{Buyer Agent} \rightarrow \text{Provider Agent} \rightarrow \text{AgentVerify} \rightarrow \text{Verification Engines} \rightarrow \text{Attestation} \rightarrow \text{Reputation} $$
The verification engines independently evaluate different aspects of the provider's delivery.
AgentVerify Service
The primary callable service is:
AgentVerify_VerifyDelivery
Input
{
"original_request": "...",
"agreed_requirements": [],
"provider_output": {},
"expected_schema": {},
"evidence_required": true
}
Output
{
"status": "PARTIAL",
"score": 78,
"confidence": 91,
"requirement_results": [],
"schema_valid": true,
"evidence_results": [],
"missing_items": [],
"risk_flags": [],
"recommended_action": "REQUEST_REVISION",
"attestation_id": "...",
"audit_receipt": {}
}
This makes AgentVerify directly usable by other autonomous agents rather than requiring a human to manually inspect every transaction.
Cryptographic Attestation
Every verification can generate a tamper-evident audit receipt.
The receipt contains information such as:
- Verification ID
- Request ID
- Provider identity
- Verification status
- Score
- Confidence
- Requirement results
- Evidence results
- Risk flags
- Timestamp
- Input hash
- Output hash
SHA-256 hashing is used to provide integrity verification.
Conceptually:
$$ H_{input} = SHA256(input) $$
$$ H_{output} = SHA256(output) $$
These hashes allow the verification record to be checked for unexpected changes.
Reputation Layer
AgentVerify does not stop at one transaction.
Verified outcomes become reputation signals for provider agents.
The system creates a feedback loop:
$$ \text{Transaction} \rightarrow \text{Verification} \rightarrow \text{Attestation} \rightarrow \text{Reputation} \rightarrow \text{Better Agent Selection} $$
Over time, agents with consistent successful deliveries can build stronger reputation signals.
Agents with repeated failures, missing requirements, or unsupported claims can be identified as higher-risk providers.
Trust Flywheel
The long-term architecture creates a trust flywheel:
Agent Transaction
|
v
AgentVerify
|
v
Verified Delivery
|
v
Reputation Signal
|
v
Better Agent Selection
|
v
More Successful Transactions
|
+------------------+
|
v
More Reputation
This creates the foundation for a more reliable autonomous agent marketplace.
SharedOS Integration
AgentVerify follows SharedOS's deny-by-default security model.
The service uses explicit capability authorization rather than trusting caller-provided permissions.
The service is restricted to the capabilities required for verification.
The intended SharedOS purpose is:
Verify and attest agent-to-agent service engagements for quality and completion.
AgentVerify does not require arbitrary:
- Shell execution
- Filesystem access
- Dynamic code execution
- Unrestricted network access
- Unnecessary system privileges
This keeps the verification service bounded and auditable.
SharedNet Integration
AgentVerify is designed as a callable SharedNet service that other agents can discover and invoke.
The service exposes:
AgentVerify_VerifyDelivery
Other agents can provide a transaction's request, requirements, delivery, schema, and evidence requirements and receive a structured verification result.
This makes AgentVerify infrastructure for the agent ecosystem rather than simply a standalone web application.
Challenges we ran into
Verifying delivery instead of response quality
The biggest challenge was recognizing that a response can be well-written while still failing the original agreement.
We therefore focused on requirement-level verification instead of relying only on general response quality.
Handling incomplete evidence
Verification becomes difficult when the provider does not provide enough information.
Instead of forcing every transaction into a binary pass/fail result, we introduced confidence and uncertainty handling.
Making the score explainable
A single unexplained AI-generated score would not provide enough trust.
We designed a transparent scoring model where every major component contributes to the final result.
Security and authorization
Because AgentVerify itself operates in an agent ecosystem, the verifier must also be trustworthy.
We therefore incorporated SharedOS capability-based authorization and deny-by-default execution.
Creating useful reputation
A reputation score should be based on actual verified transactions rather than arbitrary ratings.
AgentVerify generates reputation signals from verification outcomes and attestations.
Accomplishments that we're proud of
We are proud that AgentVerify is designed as an actual callable infrastructure service for agent-to-agent transactions.
Our major accomplishments include:
- Requirement-level delivery verification
- Completeness detection
- JSON schema verification
- Evidence verification
- Contradiction detection
- Safety and refusal handling
- Transparent weighted scoring
- Separate confidence estimation
- Uncertainty-aware verification
- Machine-readable verification verdicts
- Cryptographic attestations
- Tamper-evident audit receipts
- Agent reputation signals
- SharedOS permission enforcement
- SharedNet callable service architecture
- Multiple success and failure scenarios
- Full-stack dashboard and verification interface
Most importantly, we moved the concept of trust from a subjective rating to a verifiable transaction history.
What we learned
We learned that autonomous agents need more than the ability to communicate and execute tasks.
They need accountability.
A provider agent saying that it completed a task is not enough. The requesting agent needs evidence that the delivery actually satisfies the agreement.
We learned that reliable agent infrastructure requires:
$$ \text{Execution} + \text{Verification} + \text{Evidence} + \text{Accountability} $$
We also learned that uncertainty should be treated as a first-class outcome. When an agent does not have enough evidence to make a reliable decision, the system should say so rather than pretending to know.
Most importantly, AgentVerify changed our perspective on the agent economy:
The future of autonomous agents is not only about agents that can act. It is about agents that can be trusted.
AgentVerify provides the verification layer that can make that trust measurable.
Log in or sign up for Devpost to join the conversation.