Inspiration
Compliance gap triage is a real, recurring bottleneck I've seen firsthand in security/GRC work — a control assessment comes back with dozens of results, and none of it tells you what's actually urgent, who should own the fix, or what to tell leadership. That process — manually reading through every failure, ranking severity, writing tickets, and drafting a status update — is exactly the kind of repetitive, judgment-adjacent work the "Agents for Humans" brief asks for: take the busywork off a professional's plate, and only surface a real decision when one's actually needed.
What it does
ComplyAgent reads a compliance control dataset, classifies every gap by risk severity, drafts a remediation ticket for each of the top gaps (title, plain-language risk description, suggested owning team, priority), and produces an executive-ready summary — an overall score, the top 5 risks, and a plain-language email draft ready to send to leadership.
It isn't locked to one framework. It ships with a built-in protocol library covering five widely used standards — PCI-DSS v4.0, SOC 2 Trust Services Criteria, ISO/IEC 27001:2022, the HIPAA Security Rule, and NIST CSF 2.0 — and accepts a custom uploaded JSON dataset for anything else. The same severity engine runs unchanged across all of them.
The web UI also shows instant, deterministic metrics (pass rate, gap count, critical-gap count) the moment a protocol is selected, with zero AI calls — before the full AI-generated report is requested separately. A companion product site embeds the live app directly in a "Try It Live" section, so a visitor can run a real compliance check without ever leaving the page.
How we built it
ComplyAgent is a Strands Agents SDK agent orchestrating a four-tool pipeline:
get_control_status — loads the selected protocol dataset, groups controls by pass/fail/partial classify_risk — ranks failing/partial controls by severity, weighted by control family, with a graceful fallback for control families it doesn't recognize draft_remediation — writes a structured ticket for each top gap generate_summary — produces the overall score, top risks, and the exec email draft
The severity-scoring and ticket-drafting logic is deterministic Python, unit-tested independently of the LLM. The model's job — via Claude Sonnet 4.5 on Amazon Bedrock — is orchestrating the pipeline correctly and producing clear natural-language output, not deciding pass/fail itself. That split was a deliberate design choice: I wanted the facts auditable and testable, and the judgment/communication left to the model.
The app is built with Streamlit and deployed on Streamlit Community Cloud, using Bedrock's cross-region inference profile for Claude. The companion marketing site is a separate static build deployed on Vercel, embedding the live app via Streamlit's built-in ?embed=true mode.
Challenges we ran into
Getting Bedrock working reliably was the biggest technical hurdle. Two specific errors along the way: newer Claude models on Bedrock reject the plain model ID for on-demand invocation and require a cross-region inference profile ID (the us. prefix) instead — a ValidationException that isn't obvious until you hit it. And with a 40-control dataset, the default token budget for the agent's responses wasn't enough to carry the full pipeline through all four tool calls without truncating mid-run — raising max_tokens fixed it, but it took a MaxTokensReachedException to surface the problem.
Building solo also meant every debugging session — Git remote misconfiguration, AWS credential session handling, a merge conflict between a local README and GitHub's auto-generated one — had to get resolved without a teammate to split the load with.
Accomplishments that we're proud of
Getting a real, end-to-end agent run working — not a mocked demo, but Claude genuinely orchestrating all four tools in sequence via Bedrock, against a real (synthetic) 40-control dataset — was the milestone that made everything after it feel achievable. Extending that same, unmodified severity engine across five different compliance frameworks with zero code changes was proof the architecture actually generalizes rather than being a one-off PCI-DSS demo. And shipping a fully embedded, live-functioning product site — not just a working backend — addresses the part of the judging criteria (Design) that's easy to underinvest in when the technical build is already hard enough on its own.
What we learned
The clearest lesson: separating deterministic logic from LLM judgment isn't just a nice architectural idea, it's what makes an agent trustworthy. Unit-testing the severity scoring independent of the model meant I could be confident the agent's final report was grounded in real, verifiable numbers — not something the LLM might quietly get wrong. I also learned that AWS Bedrock's newer-model quirks (inference profiles) and token-budget planning for multi-tool pipelines are the kind of practical details that don't show up until you're actually running against real-sized data, not toy examples.
What's next for ComplyAgent
Growing the protocol library further, since the architecture already supports adding a new framework in a single line. Deploying via Amazon Bedrock AgentCore for better observability and a more production-grade runtime. And eventually, allowing ComplyAgent to pull control status directly from real scanning tools (like OpenSCAP output) rather than requiring a pre-formatted JSON dataset, closing the loop between an actual security scan and the triage report.
Built With
- amazon-bedrock
- claude
- pytest
- python
- strands-agent
- streamlit
- vercel
Log in or sign up for Devpost to join the conversation.