Inspiration

Every growing company relies on contractors and external agencies, but very few have the administrative bandwidth to manually cross-reference monthly invoices against dense Statements of Work (SOWs). As a result, rate creep, unapproved rush fees, and breached billable hour caps slip through unchecked, costing businesses thousands of dollars in silent overpayments.

When exploring generative AI for financial compliance, an immediate problem became clear: LLMs struggle with basic arithmetic. Asking a language model to calculate fractional hourly rates, sum complex line items, and deduct contract ceilings invariably leads to mathematical hallucinations.

ContractSentry was born to solve this dilemma through a clear guiding philosophy: the model reasons, typed code verifies. We set out to build an autonomous background sentinel that handles high-friction document parsing and vendor communication while enforcing strict, deterministic code guardrails to guarantee 100% mathematical accuracy.

What it does

ContractSentry is an autonomous background agent that audits every vendor invoice against its signed Statement of Work before a human ever has to look at it.

Invoices reach the agent two ways: teams can drop a raw PDF or image directly into the workspace, where Nova Pro's multimodal vision extracts the line items, or invoices can be pushed automatically from a billing system through a headless webhook (POST /api/webhook/invoice) — no manual entry required either way.

Once ingested, the agent fetches the vendor's contract terms and runs every line item through a deterministic TypeScript audit engine that checks hourly rate ceilings, monthly hours caps, and approved billing categories down to the exact cent. Clean invoices are marked APPROVED and cleared for payment. Flagged invoices get an itemized violation report, a clause-citing dispute email drafted automatically, and an instant Discord alert — but the actual dispute dispatch or payout authorization always waits for a human to click the button. ContractSentry decides what needs attention; a person still decides what happens next.

How we built it

ContractSentry is built on Next.js 16 (App Router, Turbopack) using TypeScript 7 in strict mode, orchestrated by the @strands-agents/sdk, and powered by Amazon Bedrock (Nova Pro).

1. Ingestion Layer

  • Autonomous Webhooks: A headless API route (POST /api/webhook/invoice) allows enterprise ERPs, Stripe billing triggers, or automated email parsers to dispatch invoice payloads directly to our agent daemon without human intervention.
  • Multimodal Vision: For manual reviews, users can drop raw invoice PDFs or images directly into the workspace. Amazon Nova Pro's native multimodal capabilities extract vendor data, invoice dates, and line items into a structured JSON schema validated by Zod 4.

2. The Deterministic Audit Engine

Once an invoice is ingested, the Strands Agent executes an autonomous audit sequence:

  • fetchContractRules Tool: Retrieves the active contract ceilings from the contract rules store, including hourly caps, allowable billing categories, and clause citations.
  • auditLineItems Tool: The model delegates all calculations to an isolated deterministic TypeScript engine.

The engine evaluates line items against contract boundaries using strict mathematical checks:

  • Rate Variance: Variance(rate) = Max(0, Billed Rate - Ceiling Rate) × Billed Units
  • Cap Variance: Variance(cap) = Max(0, Total Units - Monthly Cap) × Weighted Average Rate
  • Total Overcharge: Total Overcharge = ∑ Rate Variances + ∑ Unauthorized Line Items + Cap Variance

Because all arithmetic runs inside typed TypeScript callbacks (auditLineItemsDeterministic), the model cannot hallucinate currency totals or violation amounts.

3. Human-in-the-Loop Arbitration & Alerts

When discrepancies are detected, ContractSentry never takes unilateral financial action:

  • Real-Time Dispatch: A live notification is dispatched via Discord webhooks with a formatted embed detailing the vendor, overcharge amount, and violated clauses.
  • Dispute Drafting: The agent auto-generates a polite, professional dispute email citing the exact contract clause violated (e.g., SOW-2026-Section 4.2).
  • Human Decision States: Managers can click Send Dispute to dispatch the notice via Discord and mark the invoice as disputed, or select Override & Authorize to force-approve legitimate rush exceptions with a mandatory justification reason.

Challenges we ran into

  • Taming LLM Hallucinations in Financial Calculations: Initial tests revealed that prompting the model to perform the audit in natural language led to inconsistent rate variance totals. We completely extracted the math into a standalone function, auditLineItemsDeterministic(), restricting the Strands Agent's role strictly to tool invocation and contextual reasoning.
  • Fuzzy Category Matching: Real vendor invoices rarely use the exact casing or phrasing found in legal contracts. If an SOW listed "Backend Development", an invoice billing "backend development " threw false-positive unauthorized category flags. We built a case-insensitive, whitespace-tolerant string normalization pipeline to prevent false violations.
  • Balancing Headless Autonomy with Safe Controls: Designing an agent that runs in the background via webhooks while respecting the "Agents for Humans" mandate required thoughtful state machines. We ensured that while ingestion and auditing happen autonomously, outbound dispute dispatching and payout authorization strictly require human review and action.
  • Rigorous Testing Without Burning Tokens: Running live LLM inference to test every UI layout edge case would have burned Bedrock tokens unnecessarily. We created an offline, deterministic evaluation runner (src/lib/eval-suite.ts) executing 12 assertions across 4 distinct scenarios (clean invoices, case tolerance, custom SOW limits, and multi-violation invoices), proving system integrity in milliseconds.

Accomplishments that we're proud of

  • Zero hallucinated numbers, provably. Every dollar amount, rate comparison, and hours calculation runs through deterministic TypeScript, not the LLM — and we can prove it with an offline eval suite covering 12 assertions across 4 real-world scenarios, all running in milliseconds with zero Bedrock calls.
  • A full pipeline that actually runs end to end, not a mockup: webhook ingestion, contract lookup, deterministic audit, Discord alerting, and one-click dispute drafting all connect and work against a live deployment.
  • Catching a real-world data problem before it became a production bug. Real invoices don't match contract text exactly — "Backend Development" vs "backend development " would have caused false-positive violations. We built normalization for that before it ever bit a real user.
  • Holding the human-in-the-loop line under time pressure. It would have been faster to let the agent auto-send disputes or auto-authorize payouts. We kept every outbound action behind a manual click, because that was the actual point of building this for the "Agents for Humans" track.

What we learned

Building ContractSentry reinforced that the most reliable AI agents are hybrid systems. Letting language models act as reasoning engines and API orchestrators while delegating deterministic logic, math, and validation to typed code yields enterprise-grade reliability.

Working with the @strands-agents/sdk on Amazon Bedrock showed how seamless typed tool registration can be when combined with strict Zod validation schemas. By keeping humans in control of the critical decisions and leaving repetitive contract auditing to the agent, we can eliminate costly overbilling without sacrificing operational oversight.

What's next for ContractSentry

  • A real audit trail. Override justifications and dispute records currently live in front-end state for the demo. The next step is persisting every human decision to a database, so the audit trail survives a page refresh and becomes something a compliance team could actually rely on.
  • Real payout rails. "Authorise Payout" currently marks an invoice as cleared in the UI. Connecting it to an actual payment provider would let an approved invoice settle automatically instead of just being marked ready.
  • Direct ERP and accounting integrations. Beyond the generic webhook endpoint, native connectors for QuickBooks, Xero, and similar platforms would let ContractSentry plug into a finance team's existing stack with no custom integration work.
  • A multi-vendor portfolio view. Right now each audit is one invoice at a time. A dashboard tracking overcharge patterns and dispute history across every vendor a company works with would turn ContractSentry from a per-invoice checker into an ongoing vendor-risk signal.
  • Production-grade deployment on Amazon Bedrock AgentCore for better observability and scaling beyond a single demo instance.

Built With

  • agentic-ai
  • ai-agents
  • amazon-bedrock
  • amazon-web-services
  • compliance
  • deterministic-systems
  • discord
  • discord-webhooks
  • fintech
  • human-in-the-loop
  • invoice-automation
  • llm
  • multimodal-ai
  • nextjs
  • node.js
  • nova-pro
  • pdf-extraction
  • react
  • strands-agents-sdk
  • tailwindcss
  • turbopack
  • typescript
  • vercel
  • webhooks
  • zod
Share this project:

Updates

Submission history