Inspiration

I wanted to build something for the Professional Agents track that actually helped people in the corporate world — but I didn't have a clear idea at first. My dad is an investor and trader, and my cousin is a CA, so I've grown up around people who deal with contracts, filings, and fine print for a living. Around the same time, I was watching Suits, and it got me thinking about how sharp legal work actually is — catching the one clause everyone else missed. Putting those together, I landed on the idea: an agent that catches contracts quietly breaking a company's own public promises.

What it does

Covenant reads a company's public promises and the fine print of its contracts, and flags the places where the contract quietly says less than the promise.

That gap is where compliance actually breaks. A company pledges net-zero across Scope 1, 2 and 3 emissions on its website. Then a joint-venture agreement, buried in Section 7, redefines "net-zero" to mean Scope 1 and 2 only — Scope 3 quietly gone. Nobody lied outright. The language just narrowed, one clause at a time, and it takes a careful lawyer reading two documents side by side to catch it. Covenant is that careful reader, running in the background.

When it finds a real contradiction it does two things: writes a suggested redline that puts the missing commitment back, and drafts a short escalation memo to the compliance officer explaining what it found and why it matters. Then it stops. It never edits the contract. Every finding sits as "pending human approval" until a person clicks Approve or Reject — and even then, all that click does is record the decision. A human is always the one who acts.

A few things I was deliberate about:

It stays quiet on clean documents. No memo, no noise — the point of a background watcher is that it only speaks up when something's actually wrong. It knows the difference between a contract that narrows a promise and one that's simply silent about a detail. "We cover Scope 1 and 2, excluding Scope 3" is a problem. Not mentioning Scope 3 at all isn't. Getting that line right was most of the work. If it can't read a document — a scanned PDF with no text layer, say — it says so and files it for human review, instead of pretending it read an empty page and calling it clean.

It watches a folder in real time, or runs on a schedule in the cloud, and handles PDFs, Word docs and plain text.

How we built it

The core is a four-agent setup on AWS's Strands Agents SDK. Three specialists each do one job — a Reader that pulls the facts and commitments out of a document, a Reasoner that judges whether the draft contradicts the policy, and a Writer that produces the redline and the memo. The fourth is an Orchestrator, and this is the part that makes it a real multi-agent system and not three scripts in a trench coat: the Reader, Reasoner and Writer are handed to it as tools, and it decides on its own to call them in order. We didn't hardcode "read, then reason, then write" — the Orchestrator works that out because it's the logical sequence.

Around that:

The models run through a hosted inference API (an open-weights reasoning model, gpt-oss-120b). One file is the only place that knows which provider we're on, so switching is a one-line change. Underneath the AI there's a deterministic, rule-based checker. It doubles as an offline test baseline and a fallback, so the plumbing can be tested without spending an API call. A Streamlit dashboard turns raw findings into a "case file" view — matter title, the policy language and the draft language side by side, the suggested fix, and Approve/Reject. Document intake handles the formats — pypdf with pdfplumber as a backup for PDFs, python-docx for Word, including text inside tables, which is exactly where clauses love to hide. A GitHub Actions workflow runs the same scan on a timer, on a fresh cloud machine, and commits what it found back to the repo. The machine is thrown away after each run, so the repo itself is the agent's memory.

Challenges we ran into

The biggest one was self-inflicted and took me a while to see. On a realistically long contract our runs kept failing with "request too large" — but only after a few tries. It turned out our retry logic was the cause. A Strands agent remembers its conversation, and a failed call doesn't erase the attempt, so every retry re-sent everything before it: attempt two was twice the size, attempt four was four times. A comfortable 5,500-token request grew into 33,000 and got rejected outright. The retry loop was manufacturing the exact failure it was retrying against. The fix was to rewind the agent's memory before each attempt so every try is the same size — obvious in hindsight, invisible until we measured it.

The token budget was tight in general. A single orchestrated review re-sends both documents plus every intermediate result into one final call, which overruns a free tier on a real-length document. Rather than just demand a bigger budget, we added a paced fallback: the same agents, same prompts, run one at a time with pauses, so a long document still gets a genuine review instead of a crash.

Other things that fought me:

Teaching the Reasoner not to cry wolf. Early versions flagged everything, because "the draft doesn't mention Scope 3" reads a lot like "the draft excludes Scope 3" if you're careless. Silence is not a contradiction, and drawing that line precisely was real work. Honesty bugs. We kept finding places where the system said something about itself that wasn't true — a label naming the wrong provider, a case file announcing "a redline was prepared" without actually showing it. For a compliance tool that isn't cosmetic; it's the whole credibility gone. We hunted them down. The boring cross-platform stuff — the same document fingerprinting differently on Windows and Linux because of line endings, which would have made the cloud scan re-review everything on every run.

Accomplishments that we're proud of

That it never touches a contract. It would have been easy to add an "auto-fix" button and call it magic. We didn't, and we built the system so it can't — approval only records a decision, it never edits a document. Human-in-the-loop isn't a tagline here, it's enforced in the code.

That it works on domains we never tuned it for. We built it around emissions commitments, then pointed it at a data-privacy policy — "we will never sell your data" versus a draft that quietly allows licensing data to commercial partners — and it caught the narrowing with zero changes to the agents or their prompts. That told us it's reasoning about the documents, not pattern-matching on climate vocabulary.

And that it's honest about its limits. It stays silent on clean documents so it doesn't bury a reviewer in noise, it refuses to guess at a scanned PDF, and it flags a document it couldn't read rather than letting it slip through as clean. We tested it against real, published commitment language (anonymized), not just toy examples we wrote to make ourselves look good.

What we learned

Multi-agent systems aren't free. Every handoff re-sends context, and that cost stays invisible until it isn't. We learned to measure token growth, not assume it. "Just retry it" can be the bug. Our worst failure came from retry logic that made each attempt bigger than the last. Now we don't add a retry without asking what state it's dragging along. Running the thing beats testing the thing. Two of our nastiest bugs — a redline that was generated but never displayed, and a dialog that popped back after you closed it — passed every command-line test and only surfaced when we clicked through the real interface. For compliance, honesty beats cleverness. A confident wrong answer is worse than an honest "I couldn't read this," and nearly every design call came back to that. The hard part wasn't finding contradictions. It was not raising false ones. Anyone can build something that flags a lot; building something a busy lawyer would actually trust means being right about when to stay quiet.

What's next for Covenant

Add document chunking so genuinely large contracts (50+ pages) can be reviewed without hitting model limits Expand beyond emissions and privacy into more compliance domains Add proper access control and tamper-proof audit logging for real multi-user, multi-client deployment Add OCR support so scanned contracts aren't just honestly refused, but actually reviewable

Built With

  • github-actions
  • groq
  • litellm
  • multi-agent-systems
  • pdfplumber
  • pypdf
  • python
  • python-docx
  • strands-agents-sdk
  • streamlit
Share this project:

Updates

Submission history