ballot-back

A local-election ballot explainer. Type an address, get the next election, every contest on that ballot (local races first), side-by-side candidate comparisons, incumbent voting records pulled from the city's legislative system, an importance-weighted issue-alignment quiz, and a printable voter guide — with a hard rule that every factual claim links back to its source.

Working name during the build: BallotBack. Pilot city: Madison, WI.

Track

Connected Communities

Inspiration

The idea for BallotBack started in my civics class. My teacher mentioned that across the United States, roughly seventy percent of people do not vote in local elections. That number surprised me. Local elections decide many of the things that affect our everyday lives, from zoning and schools to transit and public safety, yet most people never cast a vote in them.

I didn't understand why until the next election. I asked my mom how she filled out her ballot, and she told me she left most of it blank. She recognized the names at the top, but once she reached the local races, she didn't recognize anyone and didn't know where to find information about them. So she skipped them.

That answer stuck with me. People aren't necessarily choosing to ignore their communities. They often just don't have an easy way to understand who is on their ballot or what these local offices actually do.

I wanted to change that. I wanted to make it easier for ordinary people to understand their local ballot, become more involved in their neighborhoods, and participate in the decisions that directly affect them. If more people have access to the same clear information, I hope neighborhoods can become more engaged and more closely aligned in understanding the candidates and issues they are voting on.

That's why I built BallotBack: to make local elections easier to understand and local democracy easier to participate in.


What it does

Local elections decide zoning, transit budgets, school funding, and policing budgets, yet turnout runs 15–30%. The dominant failure mode isn't apathy — it's that a voter opens the ballot and sees six names they have never heard of. Existing civic tools either cover federal races only or bury municipal data behind PDFs and paywalls. ballot-back answers three questions in under five minutes:

  1. What's on my ballot at the next election? — election date + every contest, local races first.
  2. Who are these people and what do they say they'll do? — side-by-side candidate cards with consistent columns so the eye can scan across.
  3. If they already hold office, what have they actually done? — the incumbent's legislative record, pulled from the city's Legistar system, with a plain-language summary and a source link on every item.

Shipped feature surface

Route Feature Status
/ Landing + address entry + city selector ✅
/race/[id] Race comparison view (candidate cards) ✅
/candidate/[id] Full candidate profile (positions, sources, incumbent badge) ✅
/candidate/[id]/record Incumbent voting record, topic-tag filters, disclosure banners ✅ (differentiator)
/quiz → /quiz/results 7-issue questionnaire → importance-weighted alignment scoring w/ per-issue breakdown ✅
/guide Printable personal voter guide ✅
/calendar Election dates, registration deadlines, voting info ✅
/methodology Data sources, freshness, honest gap disclosure ✅

Deliberate design constraint: neutrality-in-code

The tool is worthless if it reads as partisan, so several rules are enforced at the code level rather than left to editorial judgment:

  • Candidates are ordered by official ballot order or alphabetically — never by score, funding, or incumbency.
  • The alignment score is computed only from the user's own answers. There is no global ranking and none is ever shipped.
  • LLM summaries run under a constrained prompt that forbids evaluative adjectives; the official title is always displayed next to the summary so it can be audited.
  • Party affiliation renders as a neutral label — no color-coded emphasis implying a default.
  • Data-freshness timestamps and source links on every page.

How I built it

Stack at a glance

Layer Choice Version / notes
Language TypeScript 7.x (project-side tsconfig uses paths only — baseUrl removed)
Framework Next.js (App Router) 16.x — Server Components render DB-backed pages; Route Handlers under /api/*
UI runtime React 19.x
Styling Tailwind CSS v4 via @tailwindcss/postcss + autoprefixer + postcss
Icons lucide-react 1.x
ORM Prisma 6.x (prisma-client-js generator)
Database PostgreSQL on Neon (prod) ← SQLite (dev) provider switched mid-build; see migration story below
LLM (summaries) Google Gemini 1.5 Flash via @google/generative-ai replaced the originally planned Anthropic path
Scheduling Vercel Cron → /api/cron/* idempotent, CRON_SECRET-guarded
Hosting Vercel build = prisma generate && next build

⚠️ A note about "this is not the Next.js you know." next dev in this version writes an AGENTS.md block warning that APIs, conventions, and file structure diverge from what's in most training data. I treated the bundled docs under node_modules/next/dist/docs/ as the source of truth and worked against TypeScript 7 + React 19 semantics (e.g. tsconfig paths-only resolution, no baseUrl).

Architecture

The workload is read-heavy and mostly precomputed, which is exactly what fits Vercel's serverless model. Slow work (fetching from external APIs, LLM summarization) happens on a schedule inside cron handlers; user requests are then just database reads rendered by Server Components.

Browser
  │
Next.js App Router (Vercel)
  ├─ Server Components ── render race / candidate / record pages from Postgres
  ├─ Route Handlers (/api/*) ── quiz alignment, guide races
  └─ Cron Handlers (/api/cron/*) ── ingestion, guarded by CRON_SECRET
        ├─ civic-sync    : Google Civic → elections + contests + candidates
        ├─ legistar-sync : Legistar → council members + matters (+ seeded votes)
        └─ summarize     : Gemini → plain-language summaries of matters
  │
Postgres (Neon)  ← Prisma Client

Every cron handler validates a bearer secret before doing any work:

function validateCronSecret(request: NextRequest): boolean {
  const secret = request.headers.get('authorization')?.split(' ')[1];
  return secret === process.env.CRON_SECRET;
}

The data model

Eleven Prisma models covering the full civic graph — Election → Race → Candidate → Official → Vote, plus Matter, Issue, Position, Donation, UserPref, and Alert. Two decisions worth calling out:

  • VoteValue is a 5-state enum — YES | NO | ABSTAIN | ABSENT | RECUSED. Absent and recused are not collapsed into "no"; conflating them would misrepresent a legislator's record.
  • JSON-in-string columns (socials, sourceUrls, topicTags, watchedIssues) — a pragmatic choice made while the datasource was still SQLite, kept after the Postgres switch to avoid a schema churn during the build.

Ingestion pipeline

1. civic-sync — Google Civic Information API. Calls GET /civicinfo/v2/elections to enumerate active elections, then GET /civicinfo/v2/voterinfo?address=…&electionId=… for a set of Madison test addresses. Contests are filtered to General/Primary, upserted by (election, office), and candidates (name, party, website, email, phone, social channels) are attached. Result on our run: 9 elections, 36 races, 64 candidates (the WI state/federal primary).

2. legistar-sync — Legistar Web API (https://webapi.legistar.com/v1/madison). Uses OData query params against two endpoints:

  • GET /officerecords?$filter=OfficeRecordBodyId eq 1&$top=100 — current Common Council members (BodyId 1), filtered client-side to records whose end date is in the future.
  • GET /matters?$filter=MatterBodyId eq 1&$orderby=MatterIntroDate desc&$top=40 — the 40 most recent pieces of legislation.

Members become incumbent Candidates each linked to an Official; matters are stored with a source URL back to the public InSite page. Result: 16 real council members, 40 real matters.

3. summarize — Gemini 1.5 Flash. Idempotent: it only selects matters where plainSummary IS NULL, up to 100 per run. The prompt is deliberately locked down — summarize only from the supplied title, exactly two sentences, neutral, no evaluative adjectives — and the summary is stored alongside the source title so a human can audit it. This runs once at ingestion time, not per request, so page loads never touch an LLM.

Topic tagging — deterministic, not model-driven

Rather than spend an LLM call (and its non-determinism) on categorization, matters are tagged by keyword match against the title across eight buckets — Housing, Zoning, Transit, Budget, Public Safety, Environment, Parks, Development — falling back to General. It's cheap, auditable, and reproducible.

Issue-alignment scoring

The quiz collects a stance (FOR | AGAINST | NEUTRAL) and an importance weight per issue, then scores each candidate with a weighted mean of per-issue matches:

$$\text{alignment}(u, c) = \frac{\sum_{i \in I} w_i \cdot m(u_i, c_i)}{\sum_{i \in I} w_i} \times 100$$

where m returns 1 for agreement, 0 for disagreement, and 0.5 when the candidate is neutral or hasn't responded — a no-response is scored as partial rather than as either endorsement or opposition. The results page renders the full per-issue breakdown, never a bare percentage, precisely because a lone number reads like an endorsement.

Frontend

Server Components do the data fetching; a handful of client components carry the interactive surface — LandingPageClient, CitySelector, IssuesTable, VotingRecordCards (with topic filters), DataProvenanceModal, plus Navbar, Footer, and a CyberBackground. Mobile was a first-class target: viewport meta, sm: typographic scaling, min-h-10/11/12 touch targets, single-column→multi-column grids, and active: states for touch feedback.

The SQLite → Neon Postgres migration

The app was developed against SQLite (file:./dev.db) for zero-setup local iteration, then moved to Neon Postgres for a real Vercel deployment. Because we'd already accumulated seeded and ingested data I didn't want to regenerate, I wrote a provider-agnostic dump/restore pair:

  • scripts/export-data.js walks the models parent-before-child (so FK order is respected on import) and writes data-export.json (~6.9k lines).
  • scripts/import-data.js replays it into Postgres after flipping the Prisma datasource provider from sqlite to postgresql and applying the migration on Neon.

The duplicate-race bug and its fix

civic-sync iterates several test addresses, and the same contest appears on each one's ballot — the first cut created a fresh Race per address, producing empty duplicate cards. Fixed two ways: inline dedupe in the sync (look up an existing race by (electionId, officeName) before creating) and a cleanup script scripts/dedup-races.js that, per (electionId, officeName) group, keeps the race with the most candidates and deletes the rest, then drops any race left with zero candidates.


Challenges I ran into

1. The address-to-officeholder API no longer exists

Google retired the Civic Representatives API in April 2025. Civic's remaining voterinfo endpoint gives you contests, and divisionByAddress gives you an OCD division ID — but nothing maps an address to the person currently holding a local seat. Our designed answer was to resolve districts ourselves: geocode the address with the US Census Geocoder, load Madison's committed alder-district GeoJSON at build time, run a Turf.js point-in-polygon test to find the district, then map district → seat → current officeholder from Legistar. I committed src/data/madison-alder-districts.geojson (all 20 alder districts) and validated the geometry approach against known addresses — but did not wire it into the request path in the time available, falling back to a set of Madison test addresses and featuring the local Madison election directly. (See What's next — this is the top item.)

2. Madison's Legistar exposes members and legislation, but not roll calls

This was the single most important data-integrity finding of the build. I scanned a full Common Council meeting — 122 legislative items — and found 0 reachable per-member votes through the API. Madison passes most items by unanimous voice/consent with no recorded individual tally, so the roll-call data the "incumbent voting record" feature depends on simply isn't published.

Rather than fake it silently or drop the differentiator, I made the honesty explicit:

  • Members and matters are ingested live and are real.
  • Individual votes are deterministically seeded — a stable hash (matterId·31 + personId·17) mod 100, weighted ~78% YES to reflect that most items pass — so the record renders and demos, but every vote is labeled illustrative on the record page, the profile block, and /methodology. I never present a seeded vote as a real roll call.

The determinism matters: re-running the sync produces identical votes, so the demo is reproducible.

3. "This is not the Next.js you know"

Next 16 + React 19 + TypeScript 7 diverge from training-data conventions (App Router Route Handler signatures, tsconfig paths-only resolution with baseUrl removed, Tailwind v4's PostCSS-plugin packaging). I leaned on the version's own bundled docs rather than assumptions.

4. Duplicate races from multi-address ingestion

Covered above — the same contest across multiple sampled addresses created empty duplicate cards; solved with (electionId, officeName) dedupe plus a cleanup pass.

5. Provider swap under time pressure

Moving from SQLite to Postgres late in a build risks losing accumulated state; the export/import scripts let us switch providers without regenerating data, which would otherwise have meant re-hitting rate-limited external APIs.


Accomplishments that we're proud of

  • A working incumbent-record feature built on real municipal data — real council members and real legislation from Legistar, tagged and summarized — while being completely transparent about the one part (roll calls) the source doesn't expose. The honesty is the feature, not a caveat hidden in a footer.
  • /methodology shipped as a first-class page, not an afterthought. It names every source, the update frequency, and the known gaps. Judges ask where the data comes from; a page that answers honestly — including what's missing — is more convincing than one that claims full coverage.
  • Neutrality enforced in code: ballot-order sorting, no global ranking, constrained non-evaluative summaries, source links everywhere.
  • A genuinely idempotent pipeline — every cron job can be re-run safely (summaries skip already-summarized rows; legistar-sync resets its synthetic election cleanly; seeded votes are stable under a hash).
  • Deployed and working on Vercel + Neon with a live database, hitting the self-imposed "shipped by hour 24" rule.

What I learned

  • Local civic data is not standardized, and coverage is the real product. The decision to go one city deep (Madison) rather than nationally shallow was the decision that let the project ship with substance. The abstraction only becomes a "product" once it survives a second city.
  • Verify data availability at hour one, not hour twenty. The roll-call gap would have sunk the differentiator if we'd discovered it late. A one-hour audit of every endpoint against real addresses de-risked the entire build.
  • Precompute aggressively; keep the request path boring. Pushing all external I/O and LLM calls into scheduled, idempotent, secret-guarded cron jobs made the user-facing app a set of fast database reads — which is exactly what serverless hosting rewards.
  • For a civic tool, transparency beats completeness. Labeling seeded data as illustrative and publishing the gaps builds more trust than a slicker facade.
  • Determinism is a demo feature. Hash-seeded votes and keyword tagging mean the app looks identical every run — no surprises on stage.

Plans I made; including the ones I didn't ship

Documenting the roads not taken, and why, because the reasoning is part of the engineering.

Plan Decision Rationale
Anthropic Claude for summaries (@anthropic-ai/sdk) Dropped in favor of Gemini 1.5 Flash (@google/generative-ai) The summary task is short, constrained, and latency/cost-sensitive at ingestion scale; Flash was the pragmatic fit and consolidated us on a single Google API-key surface alongside Civic. @anthropic-ai/sdk remains in the tree as a swappable path — the summarizer is a single module, so the provider is one import away.
Drizzle ORM Dropped in favor of Prisma An empty drizzle.config.ts survives from the initial scaffold. Prisma's declarative schema, generated client, and migration tooling let us model 11 related tables and switch database providers (SQLite→Postgres) with minimal ceremony.
SQLite in production Migrated to Neon Postgres SQLite was ideal for zero-config local iteration but isn't a fit for Vercel's serverless/ephemeral filesystem. Neon's serverless Postgres matches the deployment model; the export/import scripts made the swap non-destructive.
In-request district resolution (Census geocode + Turf.js point-in-polygon) Scaffolded, not wired GeoJSON for all 20 alder districts is committed and the geometry approach was validated against known addresses, but the runtime path falls back to featured Madison races. It's the #1 "what's next" item, not a discarded idea.
FEC API / FollowTheMoney for campaign finance Rejected FEC is federal-only (out of scope for municipal races); FollowTheMoney is state-level, stale (through 2024), and doesn't cover municipal contests. The Donation model exists, but rather than show stale or empty finance data I hide the section — municipal finance has no national API and lives in per-county CSV/PDF.
OpenStates API v3 Not used Reserved for state-house races that leak onto a local ballot; unnecessary for the Madison municipal pilot.
Star-rating accountability signal Redesigned, not built Star ratings on named individuals are brigadable, invite defamation exposure, and turn a civic tool into a review site. The intended replacement is a single structured, factual signal ("I contacted this official — did you get a response?") reported as a response rate with sample size.
Candidate questionnaire portal (/candidate/claim) Deferred (stretch) Identity verification is the hard part and unsafe to do casually; the intended design is a moderation queue that displays nothing until manually approved.
Issue alerts (post-election "your rep voted X on an issue you watch") Deferred (stretch) The UserPref and Alert models are in the schema to support it; the delivery pipeline is future work. This is the feature that turns a one-time election tool into something people keep.

What's next for ballot-back

  1. Wire live district resolution — Census geocode → Turf.js point-in-polygon against the committed alder-district GeoJSON → district → seat → officeholder, so any Madison address returns its alder and its ballot rather than a featured election.
  2. Add a second city and see what breaks. The abstraction that survives two cities is the real product — a second Legistar client is a config change, but the district and finance data won't be.
  3. Real roll calls where they exist. For cities whose Legistar does publish /eventitems/{id}/votes, replace seeded votes with real tallies and gate the "illustrative" labeling on data availability per city.
  4. Ship the deferred stretch features — the response-rate accountability signal, the moderated candidate questionnaire portal, and post-election issue alerts (UserPref/Alert are already modeled).
  5. Campaign finance, done honestly — seed the pilot city's current-cycle disclosure CSV with the filing date and a source link on every figure; never promise live finance data.
  6. Partner with civic orgs (e.g. the League of Women Voters or a local civic-tech brigade) who already run candidate questionnaires and hold the relationships that get responses — the part software can't solve — and publish the ingestion pipeline as an open dataset, since the data is more valuable than the frontend.

Built With

Share this project:

Updates

Submission history