Inspiration

Most contract software focuses on getting an agreement signed.

The operational work starts after that.

A signed vendor agreement may contain uptime SLAs, renewal notice windows, service-credit remedies, compliance deadlines, and other obligations that remain active for months or years. In practice, those obligations are often tracked manually across spreadsheets, calendars, inboxes, and memory.

I built ClauseRunner for the Agents for Humans Hackathon to explore a different model:

Turn the signed contract into an active operational workflow.

Instead of another contract chatbot, ClauseRunner is designed to monitor obligations, investigate evidence, prepare remedies, and involve a human only when consequential action actually requires authority.


What it does

ClauseRunner is an autonomous post-signature contract operations agent for procurement, vendor-management, finance, and legal-operations teams.

Its core workflow is:

Contract → Obligations → Events → Evidence → Investigation → Proposed Action → Human Approval → Execution → Audit Trail

The golden demo uses a fictional SaaS agreement with a 99.9% monthly uptime SLA.

ClauseRunner evaluates fictional uptime evidence showing 99.4% availability, identifies the SLA breach, applies deterministic remedy logic, proposes a 10% service credit worth $500, and sends that proposed action to a human approval queue.

The agent can investigate and prepare the action.

It cannot execute the consequential financial action until an authorized human approves it.

That approval boundary is enforced in code, not only through prompting.


How I built it

ClauseRunner uses the Strands Agents SDK as the operational agent layer.

The Strands workflow is supported by structured tools for:

  • contract and obligation retrieval
  • evidence investigation
  • proposed actions
  • approvals
  • execution
  • audit-event creation

I deliberately kept deterministic responsibilities outside the model.

For example:

  • SLA calculations are performed in application code
  • remedy thresholds are deterministic
  • valid workflow transitions are enforced by a state machine
  • consequential actions require persisted approval
  • execution is blocked when approval is absent
  • meaningful state changes are written to the audit trail

The goal is to combine agentic reasoning with deterministic operational controls.

Live AWS architecture

ClauseRunner is deployed as a public AWS application using:

  • Amazon ECS Express Mode for the live React + FastAPI application
  • Amazon ECR for the production container image
  • Amazon DynamoDB for contracts, obligations, actions, approvals, and audit state
  • Amazon S3 for contract and evidence artifacts
  • Amazon EventBridge for scheduled obligation checks
  • Amazon CloudWatch for operational logs and observability
  • AWS IAM for scoped runtime and deployment permissions
  • GitHub Actions with AWS OIDC for remote builds and ECR deployment without long-lived AWS deployment credentials

EventBridge invokes ClauseRunner's obligation-checking endpoint on a schedule, allowing the system to monitor contract operations in the background instead of waiting for a person to open the application.

Model integration

ClauseRunner includes a Strands integration path for Amazon Bedrock using Amazon Nova 2 Lite, as well as prepared AgentCore integration.

At submission time, live Bedrock inference remains pending AWS account-level authorization, so I am not claiming it as part of the verified live execution path.

The public application currently demonstrates the full operational workflow through its deterministic Strands fallback path, while the AWS-hosted persistence, scheduling, approval, observability, and deployment infrastructure is live.


Human-controlled execution

The most important architectural principle in ClauseRunner is:

Autonomy should not mean unlimited authority.

The agent may:

  • investigate a possible breach
  • gather operational context
  • inspect evidence
  • determine which contractual remedy applies
  • prepare an action
  • explain why the action is appropriate

But actions such as service-credit claims, termination notices, and other financially or legally consequential operations require explicit human authorization.

The execution layer verifies persisted approval before allowing the action to proceed.

Without approval, execution is blocked.

This makes the human-in-the-loop boundary part of the system architecture rather than a conversational suggestion.


Challenges I ran into

One challenge was separating tasks that benefit from agentic reasoning from tasks that should remain deterministic.

It would have been easy to let the model calculate remedies or decide whether approval was necessary. Instead, I moved those responsibilities into deterministic code so the agent can remain flexible without controlling critical policy boundaries.

Deployment introduced another set of challenges.

I wanted the project to run publicly on AWS without storing long-lived AWS credentials in GitHub and without relying on a local container runtime. I therefore built a GitHub Actions deployment path using AWS OIDC, Amazon ECR, and ECS Express Mode.

I also implemented and verified real scheduled execution through EventBridge rather than presenting scheduled monitoring as a mock feature.

The remaining external challenge is AWS account-level Bedrock authorization. The Bedrock integration is implemented, but the account currently reports the model authorization state as pending. I kept the submission explicit about that limitation rather than claiming an integration that has not been verified live.


Accomplishments that I'm proud of

  • Built a working post-signature contract operations system rather than a chat interface
  • Integrated the Strands Agents SDK with operational tools and deterministic safeguards
  • Created a real human approval boundary for consequential execution
  • Deployed the application publicly on AWS
  • Persisted operational state in DynamoDB
  • Stored evidence through S3
  • Implemented live scheduled obligation checks with EventBridge
  • Verified production observability through CloudWatch
  • Built secure CI/CD using GitHub Actions and AWS OIDC
  • Created an end-to-end golden workflow from SLA evidence to proposed financial remedy and human approval

What I learned

The biggest lesson was that useful professional agents do not necessarily need maximum autonomy.

They need the right distribution of authority.

Agentic reasoning works well for investigation, context gathering, tool selection, and preparing recommendations.

Deterministic systems are better suited for calculations, state transitions, authorization, and policy enforcement.

Human judgment remains appropriate when the action has financial, legal, or external consequences.

The resulting pattern is:

Agentic investigation + deterministic policy + human authority + auditable execution.

That is the design philosophy behind ClauseRunner.


What's next for ClauseRunner

The next step is enabling the already-prepared Amazon Bedrock / Nova 2 Lite execution path as soon as AWS account authorization is cleared.

From there, the same architecture can expand beyond SLA breaches into:

  • renewal and non-renewal windows
  • vendor compliance evidence
  • insurance-certificate expirations
  • security-document obligations
  • termination rights
  • commercial credits
  • procurement follow-up workflows

The broader goal is to make signed agreements operational systems instead of passive documents.

Built With

Share this project:

Updates

Submission history