Inspiration

Small and mid-sized B2B software companies often find enterprise sales cycles abruptly halted by a single administrative bottleneck: the vendor security questionnaire. Enterprise procurement teams routinely demand answers to 50 to 300 rigorous inquiries spanning data encryption, sub-processors, incident response protocols, and multi-factor authentication.

Founders and security engineers spend between 20 and 40 person-hours per questionnaire manually hunting through disparate Confluence pages, security wikis, and historical threads to draft repetitive justifications. Generic LLM writing tools introduce catastrophic risk in this domain—they tend to smooth over gaps or hallucinate affirmative compliance policies, exposing vendors to severe legal liability and breached customer contracts.

Groundwork was engineered around a single thesis: trust is the product. Enterprise compliance requires a zero-tolerance pipeline that treats factual grounding with mathematical rigor, enforcing an independent auditor that rejects unverified claims rather than guessing.


What It Does

Groundwork is an autonomous, multi-agent compliance copilot built on the open-source Strands Agents SDK. It automates the completion of enterprise security questionnaires, vendor assessments, and RFPs through a verifiable five-stage pipeline:

  1. Ingestion & Normalization: Ingests internal policy documents (.md, .pdf, .txt) and questionnaires (.xlsx, .pdf, raw text), decomposing complex grids into discrete, atomic compliance claims.
  2. Deterministic Retrieval: Evaluates claims against indexed policy chunks using a lightweight, zero-disk BM25 and vector retrieval engine, extracting top-$k$ evidence passages.
  3. Grounded Drafting (DrafterAgent): Generates concise draft answers restricted strictly to retrieved context. If evidence is lacking, it immediately flags "Not addressed in policy".
  4. Independent Verification (VerifierAgent): Audits the draft against the raw evidence in an isolated context window. It checks for factual support, numerical contradictions (e.g., policy says 90-day rotation; draft claims 30-day), and unaddressed controls, assigning a categorical verdict (GREEN vs. RED).
  5. Human-in-the-Loop Review Dashboard: Displays verified claims alongside raw document citations and reasoning justifications, requiring human sign-off before exporting final compliance packages.

How We Built It

The entire architecture is engineered around a model-first Python pipeline powered by the Strands Agents SDK:

  • Agent Orchestration (Strands SDK): We decoupled drafting from verification using two specialized Strands agents (DrafterAgent and VerifierAgent). The verifier acts as an unyielding compliance gatekeeper with an adversarial posture toward the drafter's claims.
  • Document Ingestion & Chunking: Engineered robust parsing logic using PyPDF2 and custom chunkers to parse complex, irregular enterprise policy documents into semantically coherent fragments.
  • Dual-Layer Architecture: Built a decoupled stack featuring an asynchronous FastAPI core mounted on high-availability cloud infrastructure, backed by an interactive dashboard for rapid verification audits.
  • Deterministic Fallbacks: Programmed hard thresholding at the retrieval stage; when evidence similarity drops below baseline confidence, the orchestrator bypasses LLM synthesis entirely and triggers an immediate RED audit flag.

Architectural Decision: Why We Bypassed Native AWS Services

While enterprise deployment patterns often default to managed AWS services (such as AWS Bedrock, OpenSearch, or Kendra), we made a deliberate architectural choice to build Groundwork without direct runtime dependencies on AWS managed AI/Vector infrastructure:

  1. Zero Data Retention & Air-Gapped In-Memory Security: Enterprise security policies and SOC 2 reports contain highly sensitive, proprietary infrastructure blueprints. Ingesting these files into multi-tenant, persistent cloud indexes (like AWS OpenSearch Serverless or Kendra) introduces data residency and cross-tenant compliance liabilities. Groundwork uses a zero-disk, in-memory retrieval engine where policy embeddings and document chunks exist strictly in volatile memory during the audit session and can be wiped instantly.
  2. Sub-Second Multi-Agent Latency: Running a multi-agent verification loop requires sequential reasoning passes per question ($N$ questions $\times$ 2 agent invocations). Cloud-managed model endpoints often incur cold-start latency and multi-second round trips. By pairing the lightweight Strands Agents SDK with ultra-high-throughput LPU inference via Groq, Groundwork processes and audits full questionnaires at sub-second speeds per claim.
  3. Zero Vendor Lock-In & Portability: Compliance tooling must be deployable anywhere—from sovereign clouds to customer-owned air-gapped on-premise servers. By standardizing our Strands agents on open-source execution patterns and OpenAI-compatible endpoints, the entire pipeline remains 100% cloud-agnostic while remaining fully packageable as an OCI-compliant container for AWS ECS or Fargate if enterprise clients demand a private VPC deployment.

Challenges We Ran Into

  • Preventing "Agent Collusion": In early agent workflows, a single agent tasked with answering and critiquing its own response frequently suffered from confirmation bias. We solved this by using the Strands Agents SDK to separate context states entirely: the VerifierAgent only receives the isolated question, the candidate answer, and the ground-truth evidence chunk, preventing it from inheriting the drafter's assumptions.
  • Retrieval Schema Mismatches: Transitioning between embedding distance calculations and BM25 token-matching scores led to pipeline schema errors where missing similarity keys crashed downstream agent loops. We refactored the vector manager into a robust, defensive dictionary-extraction pattern supporting normalized scoring bounds.
  • Dynamic File Normalization: Parsing uploaded customer questionnaires across unstructured markdown and structured multi-sheet grids required building fault-tolerant document processors that gracefully handle missing tables without losing contextual hierarchy.

Accomplishments That We're Proud Of

  • True Zero-Hallucination Gatekeeping: Validating a multi-agent system that consistently catches subtle policy contradictions and safely degrades to RED when evidence is missing.
  • Clean Strands SDK Implementation: Leveraging the Strands framework to maintain minimal boilerplate, clean agent definitions, and composable execution logic.
  • Sub-Second Auditor Latency: Achieving real-time verification benchmarks that allow enterprise security teams to review 20+ policy claims in seconds rather than days.

What We Learned

High-stakes compliance automation demands radical transparency over generative creativity. Security teams do not want a black box that pretends to know everything; they want an auditable copilot that clearly shows its sources, refuses to answer when evidence is absent, and keeps human auditors firmly in control of the final submission.


What's Next for Groundwork

  • Dynamic Framework Mapping: Automatically cross-referencing ingested internal policies against standard enterprise compliance frameworks (SOC 2 Type II, ISO 27001, FedRAMP, and NIST CSF).
  • Automated Remediation Workflows: Generating explicit policy gap reports with recommended security amendments when vendor questionnaires demand controls currently missing from the knowledge base.
  • VPC-Isolated Enterprise Deployments: Providing turnkey Terraform modules to deploy the Groundwork container stack directly into private customer AWS ECS/Fargate clusters with zero outbound data egress.

Built With

  • agentcore
  • amazon-bedrock
  • amazon-web-services
  • bm25
  • chromadb
  • compliance
  • fastapi
  • generative-ai
  • gradio
  • groq
  • llama-3
  • multi-agent-system
  • pymupdf
  • python
  • rag
  • render
  • strands-agents-sdk
  • strands-sdk
  • streamlit
Share this project:

Updates

Submission history