What it does

Berkeley City Council's agenda packet for 30 June 2026 is 1,790 pages. It was published days before the vote. Inside it: thirteen separate taxes levied on a single home, a parking ordinance that changes one street, and a police surveillance policy that had been quietly rewritten since May.

QUORUM reads the whole packet, decides which few paragraphs affect one specific household, and cites the page for every claim. Then — the part that matters — it tracks those decisions across later meetings and reports what the council actually did with them.

A full run costs $0.02 on Bedrock, $0.08 on the Anthropic provider.

Inspiration

Councils publish these packets weekly, in thousands of jurisdictions. They contain the decisions that actually touch people: a rezoning next door, a bus route cut, a rate rise, a crosswalk removed. They are published with days of notice, they are functionally unreadable, and the local newsrooms that used to read them have largely gone. Public comment periods open and close with almost nobody informed enough to use them.

That is the hackathon's theme stated literally: routine, repetitive background work a person is supposed to do weekly and never does.

The deciding factor in choosing this problem was verifiability. The source documents are real, public, free and checkable. Anyone can open the same PDFs and confirm the machine was right.

What it found

In May 2026, Berkeley's Public Safety Policy Committee reviewed a police surveillance use policy and made eleven specific requests. One was simple: tell camera owners when police access their footage.

The department's response, 30 June packet, page 1394:

"the Department will include in the RFP a request for vendors to describe the feasibility of notifying camera owners each time BPD personnel access their feed. This is not a minimum requirement for vendor selection but will be evaluated as a value-added feature…"

The committee asked for a process. What appeared was a non-binding feasibility question to vendors. The provision was not deleted — it was demoted. Anyone skimming the response would conclude the request had been granted.

The council then adopted the policy as presented: Resolution 72,369-N.S., 17 speakers, recorded in the Annotated Agenda at page 19.

QUORUM also identified what was added between the two versions — six substantive new safeguards, including a First Amendment protection barring monitoring of lawful protests, and a 72-hour notification requirement if a federal agency is given camera data.

Check it yourself in 90 seconds

  1. Open the 7 May 2026 packet (176 pages) and search repositioned to capture an area — not there.
  2. Open the 30 June 2026 packet (1,790 pages) and search the same string — pages 1395, 1402, 1413, 1415.

Two public PDFs, two months apart, and a diff.

How we built it

Strands Agents SDK, with a Graph as the lifecycle coordinator — no supervisor agent layered on top of it. Conditional edges route a packet to a deep read, to an archive path, or down an error edge that halts when too much of the document is unparseable rather than guessing at it.

Three decisions shaped the design.

Deterministic wherever possible. Fetching, parsing, segmentation, rate classification, cross-meeting identity resolution, grounding checks and arithmetic are plain Python. They cost nothing and are auditable. Strands offers hooks and steering to constrain a model mid-loop; QUORUM takes the stricter path. A hook could stop the model calling a tool too often, but it could not make a tax on your home count when triage overlooked it. Removing the choice is stronger than guardrailing it. Identity in particular is never decided by a model, because identity must be explainable: the same ordinance appears as item 14 with no number in March, and as item 1, Ordinance 8,003-N.S., two weeks later. In the first-reading packet its heading is literally ORDINANCE NO. -N.S. — the decision has no identifier until it passes.

Models only where reasoning is required, routed by cost. A cheap model triages all 51 items; an expensive one reads only the few that survive. The packet extracts to 7.6M characters — roughly 1.9M tokens — so one pass through a frontier model would cost about $5.72 in input alone, against $0.02 routed.

Bounded autonomy, enforced outside agent code. A Strands interrupt pauses for human approval and survives process death — verified across two separate OS processes via FileSessionManager. Approval then goes to a Cedar policy engine (the real one, via cedarpy), which is deny-by-default.

Deployed to Amazon Bedrock AgentCore Runtime with Memory (semantic, user-preference, summarisation and episodic strategies), verified by invoking the deployed agent against a real document and confirming it returned the same answers as the local run.

Challenges we ran into

The model got the arithmetic wrong, confidently. Asked to total the tax items, it reported $0.89/sq ft against a true $1.11167 — a specific, plausible, wrong number that did not look wrong on the page. Arithmetic moved into code and is now handed to the model as established fact.

Classification, not addition, is the hard part. The agenda uses three different rate bases phrased almost identically. The largest "per square foot" rate on the page applies only to large non-profits. Counting it overstates a 1,450 sq ft household by $1,329.36 a year.

Model triage was unstable, and it nearly went unnoticed. Across five runs on identical input, recall was 1.000, 1.000, 0.667, 1.000, 0.267. One run found 4 of 15 relevant items, overlooking eleven taxes levied on the household's own home. Any single run would have looked like a working system. Rate-bearing items now bypass model triage entirely — a tax on your home affects you whether or not a model notices — which cut recall spread by 11×.

A grounding check that blocks everything is worse than none. The first version demanded a citation for every sentence, flagging "My household has no driveway" as ungrounded. The rule is now: cite what the packet says; you need not cite your own life or your own opinion.

Bedrock access was withdrawn mid-build. On 9 September, with the project working, every Bedrock model invocation began returning "Error 002: Access to Bedrock models is not allowed for this account" — all models, all regions, while the control plane kept working. AWS Support escalated it (case 178892373400420); the Bedrock team's position is that access depends on account usage history. So the model provider became configurable: the pipeline reasons in three tiers and no longer names a vendor outside one module. AgentCore Runtime and Memory still run the agent.

Accomplishments we're proud of

A measured quality claim. Precision and recall against a hand-labelled key across repeated trials — and the honest finding is the variance, not the mean. Four genuinely debatable items are excluded from scoring so the headline does not rest on judgement calls, and the label file states that the labeller also wrote the system.

A policy engine that refuses its own author. QUORUM was built from India. Filing a comment into a Berkeley meeting on an item its operator has no stake in is astroturfing — the exact harm the integrity policy exists to prevent. So the demonstration shows Cedar blocking the author: draft complete, every quote verified against the source, human approved, and refused anyway for lack of standing in the jurisdiction.

The model can recommend an action. It cannot grant itself permission to take one — including when the operator is us.

What we learned

Every place we moved work out of the model — identity resolution, arithmetic, grounding, investigation-team composition — the system got cheaper, more reliable, and easier to defend. The model's value is explaining a decision, not producing the facts underneath it.

And a single successful run proves nothing about a system whose first stage is a language model. The variance is the measurement.

What it does not do

Stating this plainly, because a demo that hides its limits is not verifiable:

  • The OCR fallback does not OCR. When more than 25% of a packet is image-only, the error edge fires and the run halts with "evidence incomplete" rather than reasoning over text it cannot verify. The June packet is only 2.3% image-only (42 of 1,790 pages), so this path has never fired on a real meeting — it is untested against a packet that would actually trigger it.
  • No comment has ever been filed. By design. The system drafts and is refused.
  • One jurisdiction, one household profile, one labeller. Berkeley, a synthetic household on a real street, and ground truth labelled by the author. Depth over breadth is the honest claim.
  • Special-meeting packets are not segmented — they number items 1a/1b.
  • Triage recall still varies between runs. Mitigated by the deterministic rate floor, not solved; ballot measures that levy a tax are not yet covered by it.
  • Scheduling is not configured. The agent runs on AgentCore Runtime and is built for unattended operation, but every run in this submission was invoked directly. It does not currently wake itself.

What's next

  • Extend the deterministic floor to ballot measures that levy a tax.
  • OCR on the error edge.
  • Special-meeting packet formats.
  • A scheduled AgentCore deployment, so the weekly run is genuinely unattended.
  • More jurisdictions.

Built With

Share this project:

Updates

Submission history