Inspiration

Local government decides zoning, bike lanes, school budgets, and policing budgets, things that hit daily life harder than most federal policy. Almost nobody reads city council agendas anyway. They're released on inconsistent schedules, written in dense jargon ("R-2 zoning density variance" instead of "apartment building near you"), and there's no way to filter for "does this affect my street" without reading everything. By the time someone notices an item matters, the public comment window has often closed.

The result is that civic participation skews toward retirees, professional lobbyists, and people with a direct financial stake, not the broader neighborhood. We wanted to build something that closes that gap quietly: an agent that does the reading so a resident or a neighborhood volunteer doesn't have to, and that only interrupts them when there's a real, time sensitive decision to make.

What it does

CivicPulse is a Strands agent that monitors a real city's public meeting agendas (San Jose, CA, via the live Legistar Web API) and does four things:

  1. Judges relevance semantically, not by keyword. A zoning item titled "Reasonable Accommodation" shares vocabulary with "affordable housing" but is legally a disability access provision under the Fair Housing Act. CivicPulse correctly separates these into distinct priorities instead of conflating them.
  2. Classifies urgency, factoring in whether a vote is imminent and whether the item sits on the consent calendar, since a consent item passes as part of a routine batch unless someone specifically requests it be pulled for discussion. That changes what "actionable" even means for that item.
  3. Summarizes relevant, time sensitive items in plain language, stripping the legal phrasing down to what's actually being decided.
  4. Drafts a neutral public comment template, only for items judged genuinely urgent, for a human to personalize and submit themselves.

Everything renders to a static, read only digest page. Nothing is ever sent, posted, or submitted automatically. The digest also carries a visible data source and staleness banner, since a cached fallback snapshot is used automatically if the live feed is ever unreachable, and the page always says which mode it's in.

How we built it

The agent runs on a single Strands agent with five tools (fetch agenda, assess relevance, classify urgency, summarize, draft comment) rather than a multi agent architecture. That was a deliberate scope decision for a short build: one well designed agent that actually works beats an unfinished swarm.

We used Strands' native structured output support so the model only supplies judgment fields (relevance, urgency, summary, draft text), keyed back to our own ground truth data for anything factual, like links, dates, and titles. That separation meant the model could never accidentally hallucinate a date or a title, only its own reasoning outputs.

For data, we pulled real agenda items from San Jose's live Legistar API rather than synthetic data, and kept a snapshot as an automatic fallback (fetch_agenda_with_fallback) so a live demo never depends on a government website's uptime. The fallback carries its own capture timestamp, and urgency judgments are explicitly anchored to that reference date rather than an assumed "today," so a stale cached run can't confidently claim a comment window is still open after it's actually closed.

The human review surface is a static, read only HTML digest grouped by urgency (Now, On your radar, Reviewed with no action needed), with a prominent "nothing is sent automatically" notice and a one line note on every consent calendar item explaining how to actually request it be pulled for discussion if a resident disagrees with the agent's judgment.

Challenges we ran into

A real urgency threshold bug. During review, we found that classify_urgency was requiring a signal of controversy in the item's own title before flagging something urgent, even for a specific, imminent, individually voted item like a real tax and fee waiver for a named development. That's backwards: this product exists specifically because agendas hide their own stakes behind dry language, so requiring the title to announce its own importance defeats the point. We restructured the prompt so a non consent item with an imminent, undeferred vote defaults to urgent on its own, and scoped the controversy check strictly to consent items, where it actually belongs.

A three times cost blowup from a schema design mistake. Early on, requiring an urgency field in the model's structured output caused it to invent an urgency judgment for every single item just to satisfy the schema, including ones it had already found irrelevant. Making that field genuinely optional, instead of just discouraging it in the prompt, dropped one real run from 70 tool calls to 41 against the same batch of live items.

A prompt constraint that could be talked around. The system prompt told the model to draft a comment only for items at "now" level urgency, but testing showed the model sometimes drafted comments for lower urgency items anyway, which visually implies urgency they don't have. We didn't trust the prompt alone for this one: a draft is now discarded at render time in code if its item's urgency isn't exactly "now," regardless of what the model produced.

An AWS Bedrock account level outage during the final stretch. Two days before the deadline, all Bedrock model calls started failing with ValidationException: Error 002: Access to Bedrock models is not allowed for this account, an account wide block with no self service fix. We filed an AWS Support case and kept working: fixes made during the outage were verified with synthetic tests where possible, and we documented plainly which fixes still need a live re-run once account access is restored, rather than claiming verification we didn't actually have.

A time boxed AgentCore deployment attempt that we deliberately stopped. The Strands scaffolding for Bedrock AgentCore worked cleanly, but the deployment path needed CDK bootstrap, ECR image push, and IAM role creation, permissions our intentionally minimal scoped IAM user didn't have and that we chose not to grant this close to a deadline purely to chase a bonus. The scaffold's entrypoint also targets a conversational, session based agent, while CivicPulse is a scheduled batch job, a structural mismatch beyond the permissions gap. Since AgentCore is a scoring bonus, not a requirement, we stopped and documented the finding instead of pushing further.

Accomplishments that we're proud of

We're proud that most of what's above came from actually running the agent against real data and reading its output critically, not from a code review that stopped at "does it run." The urgency threshold bug in particular was only visible by looking at real San Jose agenda items and asking whether the result made sense given the product's own premise, not just whether the code executed without errors.

We're also proud of the deterministic guard on comment drafting. Recognizing that a prompt instruction alone wasn't a strong enough guarantee for the one invariant that mattered most, and enforcing it in code instead, felt like the right kind of caution for a tool that touches a real, if human reviewed, civic action.

What we learned

We learned that "the model is technically working" and "the model's judgment matches what the product actually needs" are two different bars, and only real data exposes the gap between them. Synthetic test data would very likely not have surfaced the urgency threshold bug, since a synthetic test tends to be written with the same assumptions baked in as the code it's testing.

We also learned to separate what a model should output as a matter of policy from what should be enforced deterministically in code. Prompts are good for judgment. They are not a substitute for a hard guarantee when the cost of a slip is a ready to send comment draft implying urgency that isn't real.

What's next for CivicPulse

  • Complete the live re-verification of the urgency threshold fix and the templated reasoning spot check once Bedrock account access is restored.
  • Revisit Bedrock AgentCore deployment once there's time to either adapt the batch job to its session based entrypoint contract or confirm a scheduled invocation pattern is supported directly.
  • Expand beyond a single hardcoded neighborhood priority profile to a simple per user or per group configuration, still with the same three to four sharp priorities pattern rather than an open ended list.
  • Add a second real city to validate that the relevance and urgency logic generalizes beyond San Jose's specific Legistar formatting and consent calendar conventions, rather than being quietly tuned to one city's quirks.

Built With

Share this project:

Updates

Submission history