Inspiration
We wanted a way to find the small things around Cornell that rarely show up in a map search, like a bike rack on the way to class or a curb ramp near Duffield Hall. These details matter when you’re getting around campus, but finding them often means checking in person.
Many of these features are already visible in street-level photos. The missing piece is a way to search those photos by location and by what you’re looking for.
That became the idea behind Cartographer: choose an area, describe what you need, and see possible matches on a map, with the original photo attached so you can check them yourself.
What it does
Cartographer lets you search street-level imagery for physical features.
- Choose an area. Search within the current map view, draw a region, set a radius from 25 meters to 50 kilometers, or sketch a corridor along a route.
- Describe what you’re looking for. Enter a query such as “bicycle racks,” “benches,” or “curb ramps.”
- Review the matches. Results appear on a 3D globe. Clicking a marker opens the source image so you can inspect the match. You can keep multiple searches visible as separate layers and export them as GeoJSON.
Live mode retrieves imagery from Google Street View or Panoramax and uses GPT-4.1 to analyze it. We also built a sample mode with synthetic Cornell examples so we can demonstrate the interface without making paid API calls.
The results have a few important limits. Markers show where the camera was, so the object may be some distance away. Matches still need to be checked, especially for questions about accessibility. An empty search only means we didn’t find a match in the images we sampled.
How we built it
The frontend uses React 19, TypeScript, and Vite, with shadcn/ui and Tailwind for the interface. CesiumJS powers the globe. We integrated the provided Titanium renderer and its cel-shading and edge shaders, and used MapLibre to render OpenFreeMap tiles into Cesium textures.
The backend is a Flask API that runs searches asynchronously. Results appear as images are processed, and users can cancel a search while it’s running. We separated image retrieval from visual analysis so that searching the same area for a different feature can reuse images we’ve already fetched.
MongoDB stores search progress, analysis caches, and spending records. Cache keys account for changes to the query, search area, model, prompt, and sampling settings.
We also built in controls for cost and image handling:
- Local spending limits of $4 for Google and $2 for OpenAI, with funds reserved before each paid call.
- Temporary, memory-only storage for Google preview images, with limits on retention time and size.
- Structured vision responses that include a written explanation of the visible evidence and a bounding box around the match.
For development, npm run dev starts MongoDB, Vite, and Flask together. Production uses Waitress to serve the built frontend and API, with Docker Compose configured for deployment through Coolify.
What we learned
The selected area needs to stay authoritative. If someone draws a region and then mentions another place in their query, silently moving the search makes the tool unpredictable. We kept the location controls responsible for where to search and the text query responsible for what to find.
We also learned how easily sparse results can look complete. A search samples four directions per panorama within a budget of 4–48 images. That leaves gaps, and the interface needs to make those gaps clear.
Caching made repeated exploration much more practical. Users can revisit an identical search without additional API spending and reuse fetched views when looking for something else.
Image licensing also affected the architecture early on. Google’s image-handling restrictions shaped our temporary storage approach, while Panoramax required us to retain attribution and license information.
Finally, showing the evidence matters. A bounding box and an explanation won’t prevent every incorrect match, but they give users something concrete to inspect.
Challenges we faced
Placing detections on the globe. The model identifies objects within a photo, while the available geographic coordinates describe the camera’s location. We had to represent that uncertainty clearly without suggesting we knew the object’s exact position.
Tracking spending through failures. A timeout or crash can happen after a paid request has been sent. We used persistent reservations and atomic MongoDB updates so restarting the app wouldn’t reset the spending record, and disabled automatic OpenAI retries.
Working with uneven imagery coverage. Google Street View has useful coverage but comes with reuse restrictions. Panoramax provides openly licensed imagery, though coverage can be sparse. Supporting both providers, along with sample mode, required consistent behavior and clear explanations of each mode.
Keeping the interface consistent during searches. Users can cancel jobs, change areas, and add layers while results are still arriving. We had to handle those overlapping actions carefully and share work between identical requests.
Making the map comfortable to use. Much of the work went into drawing search areas, managing layers, and inspecting source photos. Those interactions needed to remain responsive on desktop and usable on mobile, alongside the Titanium globe rendering.
Built With
- cesiumjs
- flask
- google-streetview
- mongodb
- openai
- panoramax
- python
- react
- titanium
- typescript
Log in or sign up for Devpost to join the conversation.