Inspiration

Trip planning is broken across a dozen browser tabs: flights on one site, hotels on another, weather in a third — and nothing tells you your day-3 hike lands on a rainstorm until you're there. We wanted a single AI agent that plans the whole trip and owns the consequences of its plan: real bookable prices, real forecasts, feasible schedules — not plausible-sounding text.

What it does

Planora turns one sentence — "Miami to Glacier NP, October 5–9, 2 travelers, budget economy" — into a complete, day-by-day itinerary:

  • Extracts trip constraints from free text (origin, destination, dates, travelers, budget, vibe)
  • Searches live supplier APIs for flights and hotels, resolving alternate airports when they're cheaper
  • Schedules attractions weather-aware: outdoor picks on clear days, indoor backups on bad ones
  • Scores every itinerary option deterministically across cost, convenience, and feasibility
  • Re-plans around you: re-rank by cost-vs-convenience, swap any pick and the agent rebuilds the rest
  • Learns from accept/reject feedback into a preference profile
  • Ships at full parity on web (Next.js) and iOS (Expo), with SSE progress streaming while the agent plans

For the submission video, the proof plan is deliberately broader than one golden path: preload five different trip prompts through the deployed API, warm the provider/cache layer, then use Playwright to record the live demo plus a five-itinerary montage from the saved trip IDs. The samples cover beach, culture, outdoor/weather-backup, convenience-weighted, and budget-family scenarios.

How we built it

A Python 3.12 / FastAPI backend running on Google Cloud Run with Cloud SQL Postgres.

The planner pipeline is orchestrated by Google ADK (google-adk): each step runs as a BaseAgent inside a SequentialAgent, driving the shared state through the ADK Runner. The LLM layer uses the Google GenAI SDK (google-genai) directly — Gemini 3.5 Flash-Lite handles constraint extraction, Gemini 3.5 Flash handles itinerary assembly, routed to independent per-model quota pools.

The pipeline: extract brief → gather offers → enrich costs → fetch forecasts → schedule → assemble → verify.

The core principle is LLMs judge, code decides: models only narrate and classify — every price, date, ID, and link comes from verified supplier/weather data through Pydantic schemas, and an assembly fact-guard rejects any output referencing something not in its facts (with bounded retry and a deterministic fallback). SQLAlchemy async + Alembic persist trips, offers, and traces; structlog gives structured observability; LangSmith traces every node; promptfoo runs golden-dataset evals before any LLM change merges. Next.js and Expo clients share generated OpenAPI types so they can never drift from the API.

An autonomous background trip watcher runs as an async task on the server — it polls upcoming trips, checks staleness, and re-plans around weather changes without any user request, demonstrating genuine agentic autonomy.

Google Cloud, cost, and observability proof points

  • Cloud Run: separate deployed services for planora-api and planora-web, including SSE from the API to the browser.
  • Cloud SQL Postgres: durable trips, itineraries, offers, feedback, and agent traces.
  • Artifact Registry + Secret Manager: container images and production credentials managed by Google Cloud services.
  • Cost controls: Cloud Run scale-to-zero, per-provider cache TTLs, capped offer persistence, bounded LLM retries, provider timeouts, and deterministic fallbacks when model or supplier calls fail.
  • Observability: structured logs, /metrics counters/timers, persisted AgentTrace rows, and optional LangSmith traces that fail open if credentials are absent.
  • Model usage: Google GenAI SDK for Gemini, Google ADK SequentialAgent orchestration, Pydantic response schemas, and fact guards that prevent generated prices, links, dates, or IDs from reaching users.

Five itinerary samples used in the video proof

  1. Seattle → Miami family beach trip, mid-range, museums and beach time.
  2. San Francisco → New Orleans culture weekend, jazz, food, and walking tours.
  3. Chicago → Denver outdoor trip with indoor weather backups.
  4. Los Angeles → New York premium city trip weighted toward convenience.
  5. Austin → San Diego economy family trip with beaches and family attractions.

Pre-existing code disclosure

Planora was built entirely within the hackathon submission period. The repository's first commit is dated Aug 9, 2026, and development has been continuous through the submission window (Aug 30, 2026) — the backend, web client, and mobile client were all built here, with the final week focused on adopting the required Google GenAI SDK, Google ADK, and Google Cloud infrastructure (Cloud Run + Cloud SQL). No code in this repository predates the hackathon submission window.

Challenges we ran into

Free-tier Gemini taught us the hard lessons:

  • Date hallucination — the model happily planned "October" requests in September. Fixed by making dates deterministic-first: regex parses explicit months, and a cross-check discards LLM dates that contradict the user's own words.
  • Structured-output fragility — Flash-Lite truncated mid-JSON (MAX_TOKENS), returned empty itineraries, even tripped Gemini's RECITATION filter. We diagnosed each via finishReason telemetry, then constrained generation structurally: temperature control, schema-level caps, expected-day enumeration, prose clamping.
  • Per-model quota exhaustion — our daily limit died mid-demo. Since Google tracks quotas per model, we now route extraction to Flash-Lite and assembly to Flash — two independent pools, one key.
  • Scale — one hotel search returned 2,300+ offers row-by-row-persisted. Capping per search and batching writes cut plan time dramatically.
  • Guard rails earning their keep — the fact-validator caught fabricated offer IDs, invented prices injected through supplier names, and duplicate scheduling before any user saw them.

Accomplishments that we're proud of

  • A working end-to-end agent where every LLM failure degrades gracefully instead of breaking the product — plans render deterministically whenever the model misbehaves
  • Full web + iOS parity from one API contract, with types generated, never hand-edited
  • Prompt-injection safety: malicious text inside supplier names can't leak prices or instructions into itineraries
  • An honest eval story - no LLM change ships without passing the golden dataset
  • Real bookable offers and live forecasts, not mockups

What we learned

Determinism is a feature. The moments users trust an agent are exactly the moments free-text generation fails, so the winning design puts the LLM at the edges (understanding language, writing narrative) and deterministic code at the core (math, scheduling, validation). We also learned that fail-open architecture is what makes aggressive LLM usage shippable - every model outage, quota wall, or malformed response became a fallback path instead of a crash.

What's next for Planora

  • Flight booking through Duffel's order flow (offers are already cached)
  • Multi-city itineraries and ground transport between stops
  • Group trips with shared preference profiles
  • Deeper personalization as the feedback loop accumulates signal
  • Public launch: backend deployed via Render Blueprint, TestFlight distribution, and chat-surface access (WhatsApp/SMS) alongside web and iOS

Built With

Share this project:

Updates

Submission history