-
-
Quiet state — four circle members, zero connections. No pre-existing "trust" graph is drawn; edges only appear as real, independent reports.
-
A confirmed incident. Confirmation is a deterministic function counting real reports, not an AI's guess.
-
Real evidence: live domain-age lookup (RDAP), Google Safe Browsing, and SPF/DKIM/DMARC — not a canned verdict.
-
The Responder agent's output: an account-recovery checklist + a pre-filled IC3 draft, held for human approval.
-
Two unrelated incidents confirmed simultaneously — proof the engine tracks multiple real clusters, not one script.
-
The same two-incident state in Grid layout — for scanning multiple active incidents at a glance.
-
Submit Report: any circle member can flag a suspicious message, with real HTML link-extraction underneath.
-
"Someone else..." — reports aren't capped at a fixed 4 people. Real trust circles are bigger and messier.
Inspiration
This came from something that happened to us directly, not a hypothetical. A close friend's Google Workspace account was compromised and used to mass-send a fake "voicemail" notification email to their entire contact list — a convincing spoof with a "Listen to voicemail" button that led to a credential-phishing page. One of us received it, got suspicious after checking the address bar, and reached out — only to learn our friend had no idea their account had been compromised at all. They found out from confused recipients, not from any warning system of their own, and ended up filing an IC3 report alone, with no clear checklist for locking the account down and no way to know who else had been exposed.
Looking into what actually happened to their account is what turned this from "one bad email" into a design thesis. It's a specific, repeatable mechanism: the victim entered credentials once; the attacker logged in directly; an inbox rule was set to auto-delete sent mail and security-alert notifications were silenced — both ordinary account settings, weaponized, leaving no visible trace to the real owner; then the same message went out from the real, now-hijacked account to real contacts, each one now a candidate next victim, precisely because the message genuinely came from someone they trusted. Our friend was never the only casualty — they were the first node in a chain, and had no way to know it until someone downstream reported back.
That gap — between "someone in a trust circle notices something's wrong" and "the circle actually finds out and responds" — is the whole problem. It's also, we realized, structurally identical to disease outbreak detection: an index case, a contact list who were exposed, individual reports of "something's off," and the need to correlate those independent signals into a confirmed cluster before the damage compounds. The FBI's 2025 Internet Crime Report backs up how real and urgent this is: phishing/spoofing was the single most-reported crime type in 2025 (191,561 complaints), account takeover caused $359.7M in direct losses, and victims 60+ absorbed $7.75 billion in total fraud losses — up 59% from the year before.
What it does
Goal: close the detection-and-response gap for account-compromise and impersonation attacks that spread through a trust network, by turning the circle itself — family, friends, or a small team — into a pooled early-warning and response system.
Core flow:
- Report — any circle member forwards a suspicious message the moment something feels off, whether or not it's addressed to them personally.
- Correlate — a deterministic Correlation Engine checks incoming reports against recent reports from the same circle for matching sender, domain, or pattern within a time window. One report is a hunch; several independent reports converging is a confirmed cluster.
- Investigate — the Investigator agent checks real signals: domain registration age, redirect-chain destination, and authentication headers (SPF/DKIM/DMARC) — so the verdict is backed by evidence, not a guess. It also distinguishes a genuinely hijacked account from one that was merely spoofed (the named person did nothing wrong).
- Respond, differently per person — the Responder agent drafts a distinct message per exposed person based on their actual exposure state, including a middle state we specifically designed in after a real, personal experience with it: someone who entered an email/username but stopped before a password field is meaningfully different from someone who entered full credentials. Nothing sends without human approval.
- Remediate the victim — a specific, ordered, platform-specific checklist for locking the account down.
- Escalate — pre-fills the IC3 complaint fields from evidence already gathered.
How we built it
Built on the Strands Agents SDK (Python), using the Agents-as-Tools multi-agent pattern: a top-level Orchestrator routes to two specialist sub-agents — an Investigator (intake + evidence gathering, holds only analysis tools) and a Responder (drafting + remediation + escalation, holds only those tools). Deliberately kept outside either agent's reach: the Correlation Engine, a plain deterministic function the application layer calls directly — so whether an incident is "confirmed" can never be talked around by a prompt injection embedded in a malicious report. This is the concrete implementation of "AI proposes, this decides."
Model backend is pluggable via circlewatch/model_provider.py: it tries Gemini first, then Anthropic, then falls back to BedrockModel targeting Claude Sonnet if neither key is set — see "Challenges," below, for why. Evidence-gathering uses RDAP for domain-age lookups (free, no API key) and Google Safe Browsing for reputation checks. Real HTML parsing (not just displayed text) extracts links from forwarded emails and flags anchor-text/destination mismatch as its own evidence signal — the classic "says amazon.com, links elsewhere" trick.
Challenges we ran into
AWS Bedrock access, the hard way. We built and tested the agents against Bedrock first, targeting Claude Sonnet 4.6 — that was the original, and still architecturally-supported, path. Days before submission, Bedrock invocation started failing account-wide with AccessDeniedException: Your account is currently being verified — confirmed via direct boto3 probes across both Anthropic and Amazon's own Nova models, ruling out a model-specific fix. Requesting explicit model access to Claude Sonnet 4.6 through the Bedrock console hit the same wall, so it wasn't just live inference calls being blocked — the account itself couldn't be granted access to the model at all. This wasn't a credits or billing problem on our end: the account had a $100 AWS free-tier credit from ordinary new-account registration (unrelated to the hackathon's own credit offer), confirmed active and untouched the whole time — the block was a separate identity/compliance verification hold on the Bedrock service itself, and credit balance doesn't skip that queue. Redeeming the hackathon's own $50 promotional AWS credit turned into its own dead end too — the "Redeem" button on the credits page was simply disabled, and that needed a second, separate AWS Support case. Two AWS Support tickets — one for Bedrock model access, one for the broken credit redemption — both sat unassigned past the deadline. Rather than wait on a process we didn't control, we leaned on the fact that Strands is genuinely model-agnostic: model_provider.py now tries Gemini's free tier first, Anthropic second, and Bedrock last — same agents, same prompts, same tools, just a different wire to the model. It turned two external blockers into a real demonstration of the architecture's flexibility, not just a workaround.
Tuning the correlation confidence function so it escalates on real clusters without false-positiving on coincidence. A real bug surfaced during testing: httpx doesn't follow redirects by default, and our domain-age lookup service redirects to the actual registry — every single domain, including entirely legitimate ones, was silently coming back "inconclusive" until we caught and fixed it. That reinforced a design principle we'd already committed to: fail toward uncertainty, never toward "safe" — the bug was invisible specifically because it failed safely rather than lying.
Designing the human-approval step to be fast enough to matter during a live incident without ever letting either agent send or act unilaterally.
Accomplishments that we're proud of
Getting the Correlation Engine to correctly confirm a multi-report cluster from independent reporters before any single report alone would have justified acting — the core "outbreak detection" thesis, tested and passing. Restructuring from a single agent with a flat tool list into three genuinely separated agents, each structurally unable to act outside its lane — not because a prompt tells it not to, but because the tools simply aren't in its set. The evidence-gathering tool citing real, checkable facts (domain age, redirect chain, authentication results) instead of an opaque AI verdict.
What we learned
How to design a multi-agent Strands system where the highest-stakes decision is deliberately kept outside every agent's own reasoning. The practical difference between "authenticated" and "safe" — a sophisticated attacker using a genuinely compromised, properly-authenticated account passes every authentication check cleanly, so authentication alone can rule out simple spoofing but can never singlehandedly clear a sender. How phishing infrastructure actually presents itself (domain age, redirect patterns, anchor-text mismatch) well enough to build real detection signal instead of relying on a model's intuition alone.
What's next for Circle Watch: An Outbreak Tracer for Phishing
Extend the same correlation core to other attack types that spread through trust: SMS smishing waves, and AI voice-clone "distress" calls impersonating a family member — a pattern the FBI's 2025 report specifically flags as rising, with over $5M in reported losses already. Share anonymized, non-identifying threat signals across circles, so one circle's confirmed incident can pre-warn others before they're hit. Broaden remediation coaching beyond Gmail/Workspace. Direct API integration with IC3 and platform abuse-reporting endpoints, with user consent.
Built With
- amazon-bedrock
- amazon-dynamodb
- amazon-guardduty
- amazon-web-services
- anthropic-claude
- asgi
- aws-cdk
- aws-cloudtrail
- aws-cognito
- aws-secrets-manager
- aws-waf
- css3
- google-gemini
- google-safe-browsing-api
- html5
- javascript
- pytest
- python
- rdap
- rest-api
- starlette
- strands-agents-sdk
- svg
- uvicorn
Log in or sign up for Devpost to join the conversation.