Inspiration

What it does

How we built it

Challenges we ran into

Accomplishments that we're proud of

What we learned

What's next for TripAgent

Inspiration

Planning a multi-day trip is one of those "small tasks" that quietly eats a whole evening: cross-checking opening hours, transit minutes between neighborhoods, which spots can realistically share a day, what to drop when the flight lands at 3 PM. Every existing tool either dumps a generic listicle or hands you a blank map and says "good luck". And nobody trusts an AI plan anyway — models invent hours, hallucinate hotels, and quietly compress a temple visit into the minute it closed.

So we built the agent the way a human assistant would actually work: the agent works, you keep the pen. It does the grinding research in the background, and only surfaces when there is a genuine decision to make — exactly what this hackathon's Everyday Agents track asks for.

What it does

TripAgent is a conversational travel agent for East Asia backpackers (Tokyo, Kyoto, Osaka, Seoul, Beijing, Shanghai, Hong Kong…), with two doors into the same product:

  • Agent mode — describe a trip in plain language ("5 days in Tokyo, Narita in and out, Disneyland is a must"). The agent researches real spots, real transit minutes and real opening hours through tools, drafts the itinerary, pauses and asks only when there's a real decision (a spot that can't be reached before closing, an over-packed pace, a trade-off worth your call), then keeps watching the trip in the background and sends ONE notification if reality changes.
  • Classic mode — a guided form for travellers who'd rather pick every spot themselves. Same engine, same live map.

Every displayed minute comes from a deterministic planning engine with full provenance badges: catalog-verified, cached, or estimated. The model can't invent a single time.

How we built it

  • Strands Agents SDK drives the whole agent loop: 8 tools (list_cities, search_pois, spot_detail, draft_day_plan, replay_clock, transit_route, trip_budget, ask_traveller), streaming responses over FastAPI SSE into a terminal-styled web shell (Vue 3 + Leaflet), with a SummarizingConversationManager so long re-planning threads stay in budget.
  • Human-in-the-loop via a real Strands interrupt — ask_traveller pauses the agent mid-loop with concrete options; the web UI renders the decision card and the conversation resumes on the traveller's pick. The agent is instructed to reserve interrupts for genuine judgment calls, never for questions it can answer itself from the catalog.
  • Engine-authoritative architecture — the model orchestrates, but the itinerary shown to the traveller is rebuilt by a deterministic engine: geographic nearest-neighbour day grouping, capacity packing against flight windows, closing-time clipping, arrival-day anchoring. The model's echo cannot reach the screen (its structured plan is re-derived from the engine's draft before rendering).
  • Background watcher — a deterministic (zero-LLM) watcher re-replays affected days when reality changes and pushes at most one notification: "fixed, here's what changed" or a single decision. Quiet by default, like a good assistant.
  • Model-agnostic by design — production runs DeepSeek over an OpenAI-compatible endpoint; switching to Amazon Bedrock is a one-line AGENT_MODEL=bedrock change, nothing else moves. (We prepped Bedrock + AgentCore early; individual AWS accounts can't get Bedrock allowlisted inside the window, so the default is DeepSeek — the portability is the architectural point, and it's disclosed honestly in the README.)
  • Offline test discipline — a regression suite runs with zero API calls, including verdict-flip tests that replay engine decisions on both sides of their thresholds.

Challenges we ran into

  • Trust vs. capability. An LLM that plans freely produces beautiful, wrong itineraries — invented hours, closed markets at noon. Our fix was structural, not prompt-fiddling: the model proposes and narrates, the engine disposes. Getting the two layers to hand off cleanly (tool-captured draft args → deterministic re-render → provenance badges) was most of the build.
  • The closing-time problem. A spot that "fits" on paper still arrives after closing once real transit times land. The engine now replays every day against real cached legs and re-places victims across days before anything ships.
  • Streaming interrupts over SSE. Strands interrupts had to survive a streaming web transport: pause the loop, persist the interrupt, render the decision card, resume on click — without a second agent instance.

Accomplishments that we're proud of

  • A plan you can actually fly: 5 days in Tokyo with a Narita arrival/departure puts exactly one evening spot on arrival day, gives Disneyland a full day alone, keeps days district-coherent, and ends departure day at the airport — verified by an automated 5-run stability suite against the live agent.
  • The trust story is visible, not claimed: provenance badges, engine-computed times, and a decision log of every ask.
  • One engine, two complete products (chat and form), both shipped and live.

What we learned

  • "Human-in-the-loop" only works if the agent can tell a decision from a question. Most of the craft went into keeping the interrupt budget near zero: an agent that asks constantly is just a worse form.
  • Determinism is a feature you architect, not a parameter you set. Temperature 0 does not stop a model from inventing a hotel; a rebuild layer does.
  • Strands' interrupt + structured-output + tool primitives compose cleanly — the framework stays out of the way while the trust guarantees live in our engine layer.

What's next

  • Widen the catalog (more cities, rail passes, luggage-aware pacing) and let the watcher follow price/temporal alerts between booking and departure.
  • Publish the engine's feasibility replay as a standalone check the traveller can run on any imported itinerary.
  • Swap in Bedrock/AgentCore the moment individual-account allowlisting opens — the switch is one environment variable.

Built on

All code is in the public repo (MIT). The deterministic planning engine originated in my earlier open-source project (also MIT, submitted to a hackathon that ended before this submission period) and is vendored with attribution headers on every file; the Strands agent layer, conversation management, interrupts, watcher and web experience were built new for this hackathon. Full disclosure in the README's "Prior art" section.

Built With

Share this project:

Updates

Submission history