Inspiration

NYC is one of the most resource-rich cities in the world, but knowing what's actually near you right now — a library with quiet study space, a public restroom on a long walk, a healthy grocery store in a food desert, a park to decompress in — usually means digging through half a dozen different city agency pages and open data portals that were never built for someone standing on a sidewalk with a phone. We wanted asking for a city service to feel as natural as asking a friend, without losing the fact that the answer has to come from real NYC data, never a guess.

What it does

NeighborAI is a conversational NYC civic assistant. Ask it something like "I'm hungry in Manhattan", "quiet place to study", or "restroom near Central Park", and it will:

  1. Understand your intent and figure out which of five official NYC Open Data datasets you need — Libraries, Parks, Restaurants, Healthy Stores, or Public Restrooms
  2. Search only that real dataset — it never invents a location
  3. Sort results by actual distance from you using live geolocation
  4. Summarize the results in plain language and plot them on an interactive Google Map

It also handles greetings, small talk, and follow-ups like "what about Brooklyn?" without breaking the conversation or asking you to repeat yourself.

How we built it

  • Frontend: Next.js (App Router) + TypeScript + Tailwind CSS + Framer Motion, deployed on Vercel with CI/CD from GitHub
  • Data pipeline: five NYC Open Data CSV exports (Libraries, Parks, Public Restrooms, Recognized Healthy Stores, Dining locations) cleaned and normalized into typed JSON via small Node scripts — this is the deterministic "ground truth" layer
  • Search engine: a rule-based keyword/fuzzy search over that dataset — fast, deterministic, and the only thing allowed to produce actual results
  • LLM layer, added on top rather than replacing anything: an intent router (Llama 3.3 70B via Groq) classifies each message as a greeting, small talk, a new search, or a follow-up, and extracts the dataset + location as structured JSON — but it only routes to the real search engine, it never generates a location itself. A second LLM pass turns the raw search results into a natural, friendly summary
  • Geolocation: the browser's Geolocation API + a Haversine distance calculation, so "closest" is a real number, not a placeholder
  • Maps: Google Maps JavaScript API via @react-google-maps/api for interactive pins and directions

Challenges we ran into

  • Keeping the LLM "on a leash" — the core design challenge was letting the model route and narrate, while making sure official NYC Open Data always stays the single source of truth and nothing gets fabricated
  • A subtle bug where the router's internal intent key (healthyStore) leaked into the keyword search string and silently broke every healthy-store search — a reminder that even a thin AI layer needs to be tested against the deterministic system underneath it
  • The public restroom dataset has no borough field at all, so borough search for it only works when a borough name happens to appear in a facility's own name — a real NYC Open Data quirk we had to design around, not paper over
  • Making the LLM calls fail gracefully — if the model API is briefly unavailable, the app should fall back to plain keyword search instead of breaking

Accomplishments that we're proud of

  • A genuinely additive architecture — the LLM layer can be stripped out entirely and the app still works correctly as a plain NYC dataset search. The AI adds understanding, not a dependency
  • Real distance sorting from real geolocation, not a hardcoded "Nearby" label
  • A conversational experience that handles greetings, thanks, and follow-ups without ever inventing an address

What we learned

  • Treat the model as a router and narrator, not a database — the underlying open dataset has to remain non-negotiable ground truth
  • Small mismatches between how an LLM names things internally and how deterministic code parses strings can cause silent, hard-to-notice failures — worth testing explicitly at that seam
  • NYC Open Data is genuinely useful, but inconsistent across datasets (missing fields like borough) in ways that directly shape what you can build on top of it

What's next

  • Add real borough data to the public restroom dataset (e.g. via reverse geocoding from lat/long)
  • Move from straight-line distance to walking-time ETAs
  • Add open/closed-now filtering using the operational-status fields NYC Open Data already provides

Built With

Share this project:

Updates

Submission history