Inspiration

Phishing and social-engineering scams disproportionately target people who won't reliably recognize a warning sign in the moment, and who often won't tell anyone until money is already gone — frequently older relatives. We wanted something that watches quietly in the background and automatically loops in a trusted family member the moment a message looks dangerous, without requiring the protected person to report anything themselves.

What it does

ScamGuard is an Android app that monitors a protected person's SMS and Gmail in the background. Every new message is sent to an AI reasoning agent that returns a low/medium/high risk verdict with a plain-language explanation and a concrete safe next step — not just keyword matching. Every link in a message is checked against Google Safe Browsing, and sender email domains are checked for lookalike/spoofing patterns and authentication (SPF/DKIM/DMARC) results. If a message comes back medium or high risk, the protected person gets an on-device notification and a registered family member automatically receives an SMS alert.

How we built it

The Android app (Kotlin, Jetpack Compose) polls SMS and Gmail via WorkManager and posts each message to a FastAPI backend. The backend computes deterministic evidence (domain reputation, email authentication, link safety) and hands it, alongside the message, to a Strands Agents SDK agent backed by Claude on AWS Bedrock. The agent pipeline is also deployable to AWS Bedrock AgentCore Runtime as a managed runtime, with the FastAPI backend transparently routing to it when configured. The backend itself is containerized and deployed to AWS Lightsail Containers so the whole team can use the app from their own phones without anyone's laptop needing to stay on.

Challenges we ran into

We initially assumed SMS_RECEIVED broadcasts would reach any app holding

RECEIVE_SMS, per standard Android behavior. Direct testing with adb shell dumpsys activity broadcasts proved this wrong on our test device — broadcasts were delivered exclusively to the default SMS app. We rebuilt SMS capture around WorkManager polling instead of trusting the documented behavior. Tuning Bedrock request concurrency for a burst of messages (e.g. a full inbox refresh) was counter-intuitive: lowering concurrency to reduce per-call contention actually made the whole batch slower, since fewer parallel slots meant more sequential rounds through a multi-turn agent call. We had to measure actual burst behavior rather than reason about it abstractly. Deploying to AWS Lightsail Containers (unlike App Runner or ECS) has no IAM task-role mechanism, so the running container needs its own scoped credentials. We also hit a subtle, hard-to-see bug: we granted our ECR repository's pull policy to the container service's general principal ARN, not the actual image-puller role's ARN (populated only after enabling that specific feature) — every image pull failed silently until we compared the two ARNs directly. Keeping the agent resistant to prompt injection required treating message bodies, subjects, sender names, and even external search results strictly as untrusted data the agent reasons about, never as instructions it should follow.

Accomplishments that we're proud of

We measured our prompt-hardening work instead of assuming it helped: a 24-case live-Bedrock evaluation went from an 87.5% to a 95.8% match rate against expected verdicts after hardening the system prompt. Every integration in the pipeline is real, not mocked for the demo: real Google Safe Browsing results, real domain/email-authentication analysis, real background monitoring proven with force-killed-app testing, real SMS and on-device alerts, and a durable cloud backend reachable from any teammate's own physical phone

What we learned

Verify platform assumptions empirically. Documented Android behavior (broadcast

delivery) did not match what we observed on a real device, and we would have burned much more time re-discovering that later without testing it directly. Designing every credential-gated integration (Safe Browsing, Telegram, domain intelligence) to degrade gracefully instead of crashing let three people build in parallel without blocking on each other's AWS/API setup. Separating deterministic evidence computation (domain checks, authentication results) from LLM reasoning made the agent's verdicts more reliable and easier to audit than folding everything into the prompt alone.

What's next for Scam Guard - Agent based fraud detection

Expand evaluation beyond synthetic test cases with a de-identified real-message

dataset. Move off SQLite to a persistent managed database so a backend redeploy doesn't reset registered users' data. Re-enable the already-built Telegram family-alert channel as a configurable third alert option alongside SMS and on-device notifications. Broaden regional and language coverage beyond the current North-America-first assumptions in the risk-reasoning prompt.

Built With

Share this project:

Updates

Submission history