Inspiration
A small web-design agency maintains 20–100 client sites. Nobody re-checks them after launch. The first person to notice a dead tap-to-call button is usually the client — or nobody, forever. A plumber's site with broken call links loses real revenue every day it goes unnoticed.
Checking is the worst kind of professional work: repetitive (every site, every week), boring (nothing changes 95% of the time), and judgment-heavy exactly when something does change — because a 403 is not an outage, a CDN challenge page is not a missing meta tag, and a labelled multi-location phone menu is not "inconsistent contact info."
I know this because I built the underlying scanner first and watched it make every one of those mistakes on real businesses. Naive monitoring cries wolf; agencies respond by turning alerts off. The judgment is the hard part, and the judgment is what this agent automates.
What it does
Portfolio Sentinel is a Strands agent with three tools:
- scan_portfolio — audits every site with a defect auditor that reports only verifiable defects (broken
tel:links, emptymailto:, placeholder text left in production) and marks unreadable pages INCONCLUSIVE instead of guessing. The auditor is my published Apify Actor, exposed to the agent as a native MCP tool via mcp.apify.com, with a REST fallback. - diff_against_baseline — compares against the last known state and applies three judgment rules learned from real incidents, encoded as unit-tested code:
- R1 — Inconclusive is never breakage. A page nobody received proves nothing.
- R2 — Presence beats absence. A malformed
tel:link is on the page → alert now. A "missing" tag might be a challenge stub lying to the crawler → needs two confirmed sightings. - R3 — Silence is the product. No verified change → no interruption.
- surface_to_human — called ONLY when there's a real decision to make. The model writes the client-ready note: which client, what broke, how to verify it in ten seconds, suggested next step. Fixed defects get one sentence — they're billable proof of work.
How we built it
- Strands Agents SDK with the model doing the triage: the system prompt encodes the judgment rules, the tools do the deterministic work, and the agent loop decides whether a diff crosses the "real human decision" bar.
- Amazon Bedrock as the model provider (Strands default).
- MCP for tool integration — the auditor Actor is consumed through Apify's MCP server with OAuth, so no API token ever appears in config.
- The diff engine's absence-vs-presence distinction came directly from a production incident: the same URL returned 223,502 bytes → 0 findings and 11,951 bytes → 5 phantom findings minutes apart. A CDN was serving a challenge page to datacenter IPs, and the scanner faithfully audited the challenge page. You can fake an absence with a stripped response; you cannot fake a presence.
Challenges we ran into
The hardest problem was inverted: not finding defects, but not inventing them. Three separate production bugs turned out to be the same error in different clothes — treating "we couldn't see it" as "it isn't there." The fix wasn't better scraping; it was making inconclusive a first-class result and pushing that honesty all the way up through the agent's judgment rules.
Also: a render guard that was silently dead twice — two different always-false expressions, both compiled, both ran, neither threw. The regression test now exercises each guard condition in isolation.
Accomplishments we're proud of
- The judgment rules are code, not vibes —
test_baseline.pyproves R1/R2/R3 offline, no network, no tokens. - An agent designed around not talking. Most monitoring demos maximize output; this one treats the owner's attention as the scarcest resource and interrupts only for verified, actionable breakage.
- The tool it depends on is a real published product with its own hardening history, not a demo stub.
What we learned
An agent that alerts on unverified data trains its owner to ignore it. Reliability engineering for agents is mostly epistemology: what does the tool actually know, and what is it merely failing to see?
What's next
Scheduling on EventBridge/AgentCore, email delivery of alerts, and a per-client history view an agency can attach to invoices — "here's every defect we caught and fixed this quarter."
Log in or sign up for Devpost to join the conversation.