Inspiration

Every high school sits inside a community, and some communities carry far more weight than others. The Open Data Index for Schools (ODIS) from Johns Hopkins measures that weight, which it calls community "stress", for about 23,000 US public high schools. Here, "stress" means adverse neighborhood conditions (economic hardship, adult education, health, housing, and crime), not psychological stress and not a rating of the school itself. But ODIS ships as one large CSV that is hard to explore unless you already know what you are looking for. I wanted anyone, from a researcher to a school board member, to be able to see the data, ask it questions in plain English, and find out what actually drives stress where they live.

What it does

Schoolscape (https://schoolscape.alexanderyevchenko.com) is an interactive map of community stress around all 23,595 US public high schools.

  • Every ODIS measure is a layer: the five domains (Economic, Education, Health, Housing, Crime), the Composite, the Gini index, 20 indicators, and 8 context columns.
  • Zoom from states to counties to individual schools, with hover cards and a full profile for every school and area.
  • Pick two layers and the map switches to a bivariate view, while the Insight panel reports honest statistics for whatever you have selected: Spearman $\rho$ with a 95% confidence interval and $n$, at the state, county, and school level.
  • Compare two places side by side, star schools, and share any view by link.
  • Ask the map in plain English ("compare education and health in LA County and California") and it builds the view for you; when it is unsure ("poverty in Springfield"), it asks which one instead of guessing.
  • Six narrated stories walk through what the data says, from the national picture to "where a lawmaker would look first" in each region.

How I built it

1. Early exploratory data analysis

I started by profiling the raw file: 23,599 rows, one per school, and 56 columns. Missing values come in two forms, N/A and empty cells, and I charted how much was missing per row, per column, and per state. 43% of schools are complete, the average school is missing 2.4 of 58 cells, and two indicators (lead exposure risk and park access) are missing for about half of all schools. The gaps were not random: Connecticut averaged 21 missing cells per school and Puerto Rico 9, which pointed us straight at data problems I had to fix.

2. Fixes I had to apply

  • Corrupted school IDs. 19,158 of the 23,599 school IDs (81%) had been saved by a spreadsheet in scientific notation (1.00006E+11), keeping only 6 significant digits and making 2,635 values collide. I recovered the full 12-digit IDs by matching each school to the official NCES Common Core of Data 2022-23 directory on state, ZIP, and normalized name, using the surviving digits as a hard filter so an ID could never be guessed. Result: 19,154 of 19,158 recovered (99.98%), all 4,441 intact IDs verified, zero collisions; the 4 unresolvable rows were exact duplicates in the source and were dropped.
  • Connecticut. In 2022 the Census Bureau replaced Connecticut's 8 counties with 9 planning regions and renumbered its county and tract codes, so ODIS's county-level joins silently failed. I refilled Connecticut from current public sources (ACS 2019-2023 with a tract crosswalk, County Health Rankings 2025, and the CT Department of Public Health) and recomputed its scores with ODIS's own verified weights.
  • The original download is always kept byte-for-byte unchanged, and every fix is a rerunnable, hash-pinned script.

3. Additional datasets

  • NCES Common Core of Data school directory (ID recovery) and NCES school geocodes (school locations).
  • US Census cartographic boundaries, ZCTA-tract relationship files, and ACS 5-year data.
  • County Health Rankings and the Connecticut Department of Public Health (Connecticut fill).
  • EDFacts adjusted cohort graduation rates, joined on the repaired IDs, to test which conditions actually track a school outcome.
  • Every source, method, and license is cited in CITATIONS.md.

4. Statistical analysis

  • National relationships: Spearman correlations with county-clustered bootstrap intervals and Benjamini-Hochberg correction, PCA/factor analysis, and a cross-validated regression model.
  • Regional variation: for each region I ranked which domains track overall stress, using a leave-one-out composite so that a domain is never correlated with a score it is part of.
  • Regional weights: ODIS uses one equal-weight formula everywhere, so I fit per-region models of graduation rate on the five domains and built a regionally weighted stress score.

With $n = 23{,}595$, significance is cheap: any $|\rho| > 0.013$ clears $p < 0.05$. So I lead with effect size and confidence intervals, not p-values:

$$\rho = 1 - \frac{6\sum_i d_i^2}{n(n^2 - 1)}, \qquad \text{reported as } \rho \;[\text{95\% CI}],\ n$$

5. App architecture

  • Data pipeline (Python): pandas, NumPy, GeoPandas, and Shapely turn the CSV and Census boundaries into static, versioned JSON and TopoJSON (states, counties, schools, aggregates, a place gazetteer, and story presets), verified by SHA-256 and a check command that rebuilds and diffs everything.
  • Frontend: React and TypeScript on Vite, MapLibre GL for the basemap and choropleths, deck.gl for school pins, Zustand for state, D3 for scales and charts, Tailwind and shadcn/ui for the interface, and Motion for transitions.
  • Statistics in the browser: a Web Worker computes Spearman, Pearson, bootstrap intervals, and histograms live for the current selection.
  • Ask the map: the browser finds candidate places with Fuse.js, then a Vercel serverless function asks Jev (TypeSafe AI), a model that returns typed choices with calibrated confidence rather than text, so it can pick layers and places but never invent numbers. Claude Haiku 4.5 is the fallback, then an offline parser, so the bar always works.
  • Hosting: Vercel, with the whole dataset under 3 MB gzipped as static files, immutable caching, and code splitting so the first map paints in about 1.4 seconds.
  • Quality: 700+ unit tests (Vitest), Playwright end-to-end tests, and several full QA passes.
  • How I worked: I built Schoolscape with Claude Code, directing a team of AI coding agents that each took one feature or analysis, while I set the questions, reviewed the results, and made the calls. All AI use is disclosed in CITATIONS.md.

Challenges I ran into

  • Discovering that 81% of the IDs were broken, and recovering them without ever guessing.
  • Connecticut's county change, which looked like missing data but was really a broken join.
  • Circularity: a domain always "correlates" with a composite that contains it, so I had to design tests that avoid it.
  • Correlation changes with the level you look at (Crime and Education: $\rho = 0.17$ across states, $0.40$ across counties, $0.24$ across schools), so the app always says which level a number is about.
  • Graduation data is suppressed for small schools, so I measured and disclosed the coverage bias.
  • Map details that are easy to get wrong: wrapping around the world without showing Alaska twice, making pins visible over any fill, and keeping it fast.

Accomplishments that I'm proud of

  • A clean, rerunnable dataset with 99.98% of school IDs recovered, ready for anyone to join to federal school data.
  • Findings that hold up to scrutiny, for example:
    • Missing broadband is the strongest correlate of low adult education ($\rho = 0.69$).
    • Housing stress behaves unlike every other domain, and in the West it runs backwards.
    • One formula does not fit everywhere: health stress predicts graduation up to five times more strongly in the Midwest and South than in the Northeast.
  • A tool where anyone can check those findings for themselves.
  • A very clean, modern interface

What I learned

  • Data cleaning is most of the work, and the most important part to document.
  • With big samples, effect size and uncertainty matter far more than p-values.
  • The best AI features constrain the model: letting it choose, not write, made the command bar both accurate (56/56 on our test requests) and trustworthy.

What's next

  • Bring in more school outcomes (test scores, attendance) to extend the regional weighting.
  • Track ODIS over time as new versions are released.
  • Let users export any view and its statistics for their own research.

Built With

Share this project:

Updates

Submission history