Inspiration
New York City publishes exceptional climate data, but it is spread across technical datasets and dashboards. A mayor, advocate, planner, or building owner still has to answer the harder question: what should we do next, why, and what evidence would prove it worked? Carbon Copy NYC turns those records into decisions while keeping observed facts, modeled scenarios, and assumptions visibly separate.
What it does
Carbon Copy is an interactive NYC climate-policy copilot with five connected experiences:
- Action briefs: Ask a plain-English policy question and receive an audience-specific recommendation, concrete actions, success measures, limitations, and evidence IDs.
- Greener building twins: Search 14,281 multifamily disclosures by name, address, or ZIP code. Elasticsearch finds comparable properties and calculates the peer median, emissions gap, and three lower-emissions examples.
- Electric-bus scenario: Change diesel and electricity prices and watch annual energy spending update. A Mistral-powered stress test asks “what would change your mind?” and exposes the exact break-even thresholds.
- The next 100 trees: Adjust heat vulnerability, vegetation deficit, and limited A/C access, or describe priorities in natural language. Mistral proposes explicit weights, Elasticsearch reranks neighborhoods, and the interface highlights who moved up or down.
- Climate evidence: Explore measured flooding, future flood exposure, green infrastructure, neighborhood heat indicators, and air quality.
How we built it
Elasticsearch is the evidence and ranking layer. It powers fuzzy, boosted building autocomplete; structured filters for comparable buildings; numeric sorting and peer analysis; climate-signal retrieval; and a script-score query that normalizes and weights neighborhood heat indicators.
Mistral Large 4 is the reasoning layer. It converts verified evidence into structured policy briefs, translates natural-language planning preferences into explicit tree weights, and chooses the most decision-relevant bus assumption to stress-test. Responses use JSON structured outputs validated with Pydantic.
Python owns the numbers. It calculates building benchmark gaps, tree allocations, fleet energy costs, emissions, and break-even curves. Generated numeric claims are checked against the evidence packet before they are displayed. FastAPI exposes this logic to a polished React interface; the original Streamlit Evidence Lab remains available for technical inspection.
The design creates a clear chain:
- Elasticsearch finds and ranks the relevant evidence.
- Python computes reproducible scenarios.
- Mistral selects and explains the decision-relevant story.
- Python validates the claims.
- React shows the recommendation, evidence, assumptions, and limits together.
Creative use of Mistral
Mistral does more than summarize. In the tree planner, the model changes the dashboard itself: a sentence becomes explicit weights, those weights become an Elasticsearch ranking, and the new ranking is compared with the prior plan. In the bus tool, Mistral acts as a skeptical reviewer, selecting the assumption most likely to overturn a recommendation; deterministic code then finds the break-even point and draws the sensitivity curve.
Challenges we ran into
The hardest problem was preventing a fluent model answer from becoming an unsupported policy claim. We solved that by giving Mistral bounded evidence IDs, using structured responses, performing calculations outside the model, and validating generated numbers. We also had to make very different datasets comparable without implying causation, handle public API rate limits during a live demo, and explain climate units to a nontechnical decision-maker.
Accomplishments we are proud of
- Indexed and searched 14,281 NYC multifamily disclosure records.
- Combined building, transit, heat, flood, green-infrastructure, fuel-price, and air-quality evidence in one product.
- Built interactive, falsifiable recommendations instead of a static chatbot response.
- Made every major model assumption visible and editable.
- Added explicit caveats for planning scenarios, operational emissions, and incomplete capital costs.
What we learned
The strongest role for an LLM in civic analytics is not arithmetic. It is translating human intent into transparent analytical choices, selecting relevant evidence, explaining tradeoffs, and identifying what evidence is still missing. Elasticsearch and deterministic calculations give that reasoning a trustworthy foundation.
What's next
We would add route- and depot-level transit data, block-level tree-site feasibility, a full electric-bus total-cost model, richer geospatial maps, and building retrofit records that can move the greener-twin feature from benchmarking toward verified implementation guidance.
Built With
- elasticsearch
- fastapi
- mistral-ai
- mistral-large-4
- nyc-open-data
- python
- react
- streamlit
- vite
Log in or sign up for Devpost to join the conversation.