Inspiration

AI agents can execute tasks, but execution is not proof of fulfillment. An agent can report success while the real-world commitment remains incomplete, overdue, or contradicted by evidence.

Covenant was built around one question:

What evidence proves that an AI-handled commitment was actually fulfilled?

We wanted to move autonomous systems from “done” to “proven.”

What it does

Covenant is an autonomous commitment intelligence and fulfillment verification system.

It discovers commitments, gathers and corroborates evidence, assesses risk, proposes remediation, manages approvals, executes authorized actions, and independently verifies outcomes.

Its core principle is:

$$ \text{Execution success} \neq \text{Fulfillment} $$

Agents can reason and propose actions, but Covenant's deterministic runtime controls authorization, policy, approvals, execution, and state transitions.

A commitment reaches:

$$ \texttt{RESOLVED(VERIFIED)} $$

only after fresh external evidence passes the \texttt{VerificationGate}.

How we built it

Covenant uses Strands Agents as a real multi-agent reasoning layer, not as a thin chatbot wrapper. Specialized Strands agents handle commitment reasoning, evidence synthesis, risk analysis, remediation planning, supervision, and post-execution verification. Agents produce structured decisions and invoke Covenant tools to interact with external systems.

We deliberately separated agent intelligence from runtime authority:

$$ \text{Strands Agents} \rightarrow \text{Deterministic Covenant Runtime} \rightarrow \text{Authorized Tools} \rightarrow \text{External Systems} $$

The deterministic runtime owns the correctness and governance boundaries: authorization, policy evaluation, human approvals, idempotency, concurrency control, persistent state, and lifecycle transitions.

The critical component is the deterministic \texttt{VerificationGate}. After an action executes, Covenant does not trust the agent's success message. It obtains fresh external evidence, evaluates the verification result, and only then permits:

$$ \texttt{VERIFYING} \rightarrow \texttt{RESOLVED(VERIFIED)} $$

This creates a hard boundary between what the agent reasons, what the system executes, and what the world proves actually happened.

Challenges we ran into

The hardest challenge was building a genuinely autonomous Strands system without allowing the model to become the final authority over its own success.

A capable agent can reason about a commitment, select tools, and propose remediation. But allowing its output to directly authorize actions or declare fulfillment would undermine trust.

We therefore engineered deterministic boundaries around the agent loop:

$$ \text{Reason} \rightarrow \text{Propose} \rightarrow \text{Authorize} \rightarrow \text{Execute} \rightarrow \text{Re-evaluate} \rightarrow \text{Verify} $$

Another challenge was evidence semantics. Covenant must distinguish a genuine contradiction from incomplete corroboration or a simple evidence gap, rather than treating every disagreement as a conflict.

We also had to make verification real rather than demonstrative. Persistent state, idempotent execution, approval boundaries, and fresh external evidence ensure that a successful tool call can never by itself produce a false \texttt{RESOLVED} state.

The resulting architecture is deliberately asymmetric:

$$ \boxed{\text{Strands provides intelligence and autonomy; deterministic infrastructure provides authority and proof.}} $$

Accomplishments that we're proud of

We built an end-to-end commitment lifecycle:

$$ \text{Discover} \rightarrow \text{Corroborate} \rightarrow \text{Assess} \rightarrow \text{Plan} \rightarrow \text{Approve} \rightarrow \text{Execute} \rightarrow \text{Verify} \rightarrow \text{Resolve} $$

We are especially proud of the separation between:

$$ \text{Agent belief} \neq \text{Runtime authority} \neq \text{Execution result} \neq \text{Verified outcome} $$

Covenant also provides an auditable decision trace, bounded autonomous monitoring, evidence intelligence, approval controls, concurrency protection, and a real post-execution verification loop.

The key achievement is:

$$ \boxed{\text{An agent cannot declare success without evidence.}} $$ What we learned

We learned that trustworthy agentic systems require governance around intelligence, not intelligence alone.

Models are excellent at reasoning and proposing actions. Deterministic infrastructure should enforce authorization, policy, state, idempotency, and verification.

We also learned that evidence quality matters as much as model reasoning. Explicitly separating facts, inferences, corroboration, gaps, and contradictions makes autonomous decisions more reliable and auditable.

Most importantly:

$$ \text{Autonomy without verification is only a claim of success.} $$ What's next for Covenant

Next, Covenant will expand from individual commitments into a broader verification and governance layer for autonomous workflows.

We plan to add more evidence connectors, enterprise integrations, richer long-running monitoring, stronger policy controls, and broader model-provider support.

The long-term vision is:

$$ \text{Autonomous agents act} \quad+\quad \text{Covenant independently governs and verifies} $$

so autonomous systems are accountable not for what they say they accomplished, but for what the evidence proves actually happened.

Built With

  • agent-governance
  • agentic-ai
  • ai-agents
  • ai-safety
  • amazon-bedrock
  • amazon-web-services
  • autonomous-agents
  • autonomous-workflows
  • decision-intelligence
  • enterprise-ai
  • evidence-based-ai
  • fastapi
  • gemini
  • generative-ai
  • human-in-the-loop
  • llm
  • multi-agent-systems
  • ollama
  • python
  • react
  • responsible-ai
  • strands-agents
  • typescript
  • verification
  • workflow-automation
Share this project:

Updates

Submission history