Inspiration

A few months ago I built Frenchly, an AI-powered French exam prep app for Express Entry candidates trying to earn CRS points for Canadian immigration. Once it was live, I ran into a problem every solo founder hits: I had no idea who my real competitors actually were, and even when I found a few, I didn't know what specifically I should change on my own site in response. Was a competitor's price actually a threat, or a completely different product wearing similar marketing language? Was their "new" feature actually new, or had it always been there and I just hadn't looked closely enough?

I looked at competitive-intelligence tools like SEMrush, Klue, and Crayon — they're built and priced for teams with a dedicated analyst, not a solo founder. So I built the analyst myself, as an agent.

What it does

Competitor Watch Agent takes a business profile and does the actual work a competitive-intelligence analyst would do, end to end:

  • Discovers your real competitors if you don't know them yet, using three fallback search strategies (a real "alternatives" page, a category/feature search for niches too small to have one, and the model's own market knowledge as a last resort), with every candidate independently verified as legitimate, genuinely relevant (same core function, not just shared buzzwords), and region-appropriate.
  • Grounds itself in your actual live site, not just what you typed about your business.
  • Fetches and diffs competitor pages against the last snapshot, correctly telling a first-ever look at a page apart from a real detected change.
  • Judges materiality — a typo fix vs. a genuine price change — from meaning, not keyword matching.
  • Analyzes impact specifically for your business, referencing your own pricing/positioning by name.
  • Prioritizes across every competitor into a short, honest action list (never padded to hit a quota).
  • Recommends a concrete action and names a real tool for each item, treating pricing/positioning changes as decisions to evaluate, not one-week tasks.
  • Builds a full profile of each competitor — a visual/UX critique, company background from their own public About page, and how often they actually update.

How I built it

I used the Strands Agents SDK with Anthropic Claude (Sonnet 4.5) as the model provider — no AWS Bedrock, no billed AWS service, since I only have a free AWS Builder ID. The system is one orchestrating Strands Agent that decides its own next move, calling out to a toolbox where several tools are themselves a separate nested Agent with a focused job and forced structured output: materiality judgment, impact analysis, prioritization, recommendation, the competitor-discovery pipeline, the visual critique, and a "worthiness gate" that decides whether a competitor is even legitimate and relevant before it's allowed into the pipeline at all. A single run makes 10–45+ real model calls, each doing one narrow thing well.

I built it in stages — scaffold, fetch, snapshot/diff, materiality, impact, prioritization, recommendation, orchestration — testing each stage in isolation before wiring it together, then layered on competitor discovery, a Streamlit frontend, visual UI critique via Playwright, and company-profile enrichment.

Challenges I ran into

The real challenges only showed up once I started testing against real businesses — my own Frenchly, then a real eyewear retailer, a real B2B SaaS, a real proptech company, and a real digital-health company — and reading the actual generated reports line by line instead of trusting that they sounded right:

  • A fabricated event. On a competitor's first-ever page fetch (no prior snapshot to compare against), the agent once wrote "Competitor X has pivoted dramatically" — there was no evidence of any pivot, just a first look. I had to make the first-look/real-change distinction structural: a different system prompt, a regex check that catches change-implying language and forces a rewrite, and a boolean flag that gets passed through mechanically rather than trusted to survive a chain of LLM calls.
  • Visual scores with zero variance. Scoring competitor screenshots independently made every site converge on the same "7/9, feels corporate and cold" verdict — a score with no variance is useless. Fixing it meant comparing every screenshot in one call and forcing a strict ranking, and even then it regressed once until I removed an escape hatch in the prompt and added a mechanical duplicate-score detector that forces a retry.
  • A rejected competitor still reaching the final recommendation. The worst bug: my own worthiness gate correctly flagged a competitor as not relevant, and the final report still recommended restructuring pricing to counter it — because the gate's verdict was reported next to the pipeline instead of actually filtering it. I fixed this the way I fixed the first bug: a hard filter in plain Python that drops anything not explicitly marked relevant, before the model ever sees it, so no amount of prompt drift can leak it through again.
  • Discovery blind spots. For a small, real niche (digital play-therapy tools) there was no "X alternatives" page anywhere on the internet, so my search-based discovery came back empty and fell through to a much weaker memory-only guess. I added a second search mode that searches by product category instead of company name and treats the organic search results themselves as candidate competitors.
  • A pattern I didn't expect: reckless pricing advice. Across multiple test businesses, the agent kept recommending pricing or packaging changes framed as "ship this week" tasks — which is exactly the kind of hard-to-reverse, high-blast-radius decision that shouldn't be a one-week task. I added a required field to every recommendation that forces pricing/positioning changes into "decision to evaluate" framing instead of "ship it" framing.

Accomplishments that I'm proud of

Every one of those bugs was caught by actually reading full generated reports against real businesses, not by assuming the pipeline worked because the code ran. I'm proud that the fixes are structural — mechanical filters and regex checks that don't rely on hoping the model behaves — rather than just rewording a prompt and hoping it doesn't regress next time.

What I learned

The gap between "the agent produces a plausible-sounding report" and "the agent produces a correct report" is enormous, and it only closes by adversarially reading the actual output against real, checkable facts — not by trusting that a well-designed prompt will hold up under real-world variance. Several of the worst bugs (the fabricated event, the leaked rejected competitor) would have looked completely fine on a skim.

What's next

Reading third-party comparison sources (not just a competitor's own homepage) to catch cases where two products share marketing language but aren't real substitutes; deploying to AWS Bedrock AgentCore; and a scheduled/recurring run mode so the update-cadence tracking has real history to work with from day one.

Built With

  • anthropic-claude-(sonnet-4.5
  • beautiful-soup
  • direct-api-?-no-aws-bedrock-needed)
  • playwright
  • sqlite
  • strands-agents-sdk
  • streamlit
Share this project:

Updates

Submission history