Inspiration
The CFPB reports that hundreds of US counties run a local elder-fraud response network — adult protective services, police, a partner bank, an area agency on aging, usually coordinated by one part-time volunteer with about four hours a week. US adults 60+ reported $2.4B in fraud losses to the FTC in 2024 alone, and the FTC's own modelling puts the real figure between $10.1B and $81.5B once under-reporting is accounted for.
The part that stuck with us wasn't the total. It was the shape of it: a crew working a neighbourhood doesn't stop after one victim. The evidence that would identify them is already sitting in a coordinator's inbox, spread across reports filed days apart by people who've never spoken to each other. Report 7 and report 19 and report 44 each look like an isolated, embarrassing phone call. Nobody has the hours to notice they're the same crew.
That's not a detection problem. It's an attention problem, at exactly the scale a background agent is good at — and it fit squarely into the Good Neighbor Agents track: an agent that helps a group of people, not one.
What it does
A county elder-fraud response network is usually coordinated by one part-time volunteer with about four hours a week. Reports arrive one at a time. Each is handled, filed, and forgotten. Nobody has the hours to notice that report 7, report 19 and report 44 describe the same crew working the same postcode.
Porchlight runs in the background and does that noticing.
While nobody is watching, it takes reports from a webhook, deduplicates repeat calls from the same resident, extracts indicators, runs allow-listed corroboration tools, merges follow-ups into existing cases, and correlates across people. When reports from different residents turn out to share a crew, it drafts a community warning and stops.
The coordinator opens the inbox and sees three things, in this order:
- what Porchlight handled while they were away;
- one prepared decision, carrying the evidence that produced it;
- the case list, ordered by hours until the money cannot be recovered.
They approve, edit, or reject. That is the whole interaction.
Not the person being scammed — that would be a personal assistant. The user is the coordinator of a local elder-fraud prevention and response network, and their work — intake, corroboration, cross-agency notification, community education — is exactly the repetitive, judgment-heavy load this hackathon asks an agent to lift off a human.
How we built it
- Strands Agents for the four judgment steps: intake (what happened, and what evidence is missing), corroboration (which allow-listed tool would settle the gap), stage assessment (a swarm of three assessors arguing about how close this resident is to irreversible loss), and campaign judgment (is this cluster really one crew).
- Deterministic correlation, deliberately not a model decision. Clustering
is union-find over hard indicators in
src/porchlight/correlation.py. It is reproducible, testable, and — the part that matters — prompt-injected prose cannot argue its way into or out of a cluster, because set membership is computed from indicators rather than from text. - A Cedar policy layer outside the model, default-deny, every decision
audited with who asked, which rule answered, and about which report — now
also loaded and
ACTIVEagainst a real AgentCore Gateway (PorchlightGateway), with the policy engine attached inENFORCEmode. - AgentCore Memory backing every agent role's conversation session,
verified live with an independent
list_eventscheck. - Deployed to AgentCore Runtime — a real
agentcore invokeagainst the running endpoint returned a correctly triaged case, processed end to end on Amazon Nova Pro inside the deployed ARM64 container. - SQLite for the queue, cases, approvals and audit trail, so the agent survives a restart and a background worker and the web process cannot disagree.
- A capability-based approval model. An approval is bound to one action, one audience, and the exact message digest, single-use and short-lived. Editing a draft invalidates it by construction.
Challenges we ran into
Four worth naming, because each was a real defect rather than a war story:
The evaluation was scoring the wrong thing. It measured only the campaign it was hoping to find, which rewards a system that clusters everything. Adding link-level scoring against crew ground truth, plus a four-case challenge set, dropped link precision from 1.00 to 0.50 and exposed 21 false links. The cause: two reports sharing one indicator were treated as the same crew — so four unrelated scams that each told the victim to "ring your bank on the number on your card" were fused into one fictitious campaign. A link now needs two independent shared indicators, or one plus the same script. Precision returned to 1.00 with zero false links.
Live mode was broken behind a green test suite — twice. Offline mode never
exercised the live path at all. First: a tool-bearing agent's structured-output
call spent its tool-use turn on the actual tools and then failed — fixed with
a two-phase call, converse first, then ask for the schema. Second, found only
once Anthropic's model was replaced with Amazon Nova Pro (see below) and a run
finally reached the stage-assessment swarm: pipeline._extract_structured()
referenced a SwarmResult.execution_order attribute that does not exist in
the installed Strands SDK — offline mode's stub never called that function, so
it was wrong behind 200 passing tests until a live run actually reached it.
Anthropic's own model turned out to be blocked by AWS billing, not by
anything in our control. AccessDeniedException: INVALID_PAYMENT_INSTRUMENT
— confirmed, via the AWS Marketplace agreement record itself, as an account
created and terminated 15 seconds later over a missing credit card. Nine other
Bedrock providers tested clean on the same account. Swapping
PORCHLIGHT_MODEL_ID to Nova Pro — one environment variable, same
architecture, same policy layer — got a full five-node pipeline running live,
end to end, with zero errors.
The demo was loading the answer. The dashboard's replay button loaded the whole corpus, so the campaign was on screen before the coordinator did anything. The replay endpoint now holds every crew one report short of threshold, read from the corpus's own ground truth, so the campaign is found live, on camera.
Accomplishments that we're proud of
Every number in the README is either produced by a command in the repo or cited to a source. The ones that were neither were deleted — including a ~254,000× "speedup" computed by dividing an unmeasured manual baseline by the offline stub's own latency.
The evaluation reports the number that makes the system look worst: held-out injection recall of 0.14–0.57 (median 0.29), quoted instead of the 1.00 we score on payloads written against our own signature list.
And past the eval harness: a real AgentCore Gateway, a real Policy engine with
all five Cedar rules ACTIVE against it, real AgentCore Memory sessions
verified independently, and a real deployment to AgentCore Runtime — verified
not by a status page, but by a live invocation that came back with a correctly
triaged case.
What we learned
A benchmark that only contains the case you designed for tells you nothing. The challenge set took an afternoon and found a bug that would have broadcast a false campaign to a list of frightened people on the strength of a bank's own customer-service line.
And offline-mode green tests can hide real bugs in the live path indefinitely — two of them, in our case — because "the tests pass" and "the live pipeline runs" turned out to be different claims until we made ourselves prove the second one, not just the first.
What's next for Porchlight
- Route Porchlight's actual tool calls through the AgentCore Gateway we
already built — the five Cedar rules are loaded and
ACTIVEthere, but the running app still enforces every decision through the in-process module. Closing that gap is the difference between a strong control and an enforced one. - Extend AgentCore Memory past conversation state into the structured record store itself, once the access-pattern mismatch (synchronous reads across a whole community vs. a conversational memory API) has a clean answer.
- Put an IdP in front of approvals so the approver is authenticated, not just named.
- Run the manual-baseline study in
docs/manual-baseline.mdwith real coordinators, and report time and accuracy.
Built With
- amazon-bedrock
- amazon-bedrock-agentcore
- amazon-bedrock-agentcore-gateway
- amazon-bedrock-agentcore-memory
- amazon-bedrock-agentcore-policy
- amazon-bedrock-agentcore-runtime
- amazon-cloudwatch
- amazon-ecr
- amazon-nova
- aws-codebuild
- aws-iam
- aws-lambda
- cedar
- css
- docker
- fastapi
- github-action
- html
- javascript
- pydantic
- pytest
- python
- sqlite
- strands-agents
- uvicorn


Log in or sign up for Devpost to join the conversation.