Inspiration

When you have a vague product idea, your first question is "has someone already built this?" But that question is unanswerable until you understand what the market actually looks like — and by then, your original assumption may not even be the interesting one.

Most market-research tools confirm what you already believe. We wanted an agent that could challenge the assumption instead, without quietly deciding your direction for you.

What it does

MarketCompass turns a fuzzy product idea into an evidence-backed position through collaborative exploration.

You start with one sentence ("I want to build an AI speaking-practice app, but not for children"). The agent confirms an anchor — hard constraints, a starting hypothesis, and a north star — then autonomously scans the market, discovers the dimensions that actually separate competitors, and teaches you the structure of the industry before asking you to choose a position.

The critical design decision: the agent has execution authority, not directional authority. When it finds something that would change your direction, it surfaces the decision back to you instead of acting on it.

The final output is a market map with your exploration path drawn on it — not just where the competitors are, but where you started, what changed your mind, and where you landed.

How we built it

Architecture: three layers with separated powers

  • Flow Control Layer — a custom ADK BaseAgent orchestrator that judges only phase transitions across six phases
  • Market Intelligence Layer — Research and Analyst agents in a conditional loop, plus an Evidence Validator that checks whether the Analyst's classifications are actually supported by retrieved data
  • User Interaction Layer — the sole user-facing exit, and the sole writer to long-term memory ## Challenges we ran into Submission-day model capacity. Gemini service degraded roughly an order of magnitude during the final hours. The same minimal prompt measured 8.1s, then 27.1s, then 62.0s within a two-hour window. Full research calls with schema timed out at 150s across every candidate model; gemini-2.5-flash returned 404 (retired for new users).

We built a model fallback chain that distinguishes capacity failures (429/5xx/timeout → degrade to next model) from real bugs (404, schema errors → raise normally). Critically, we removed the retry path: when a tool raised TimeoutError, ADK returned the error to the model, which retried the same call, burning the entire request budget in stacked timeouts.

What the system did when it couldn't research. It did not invent a market map. It logged an audit event, told the user this was a live-service failure rather than a research conclusion, and returned control. That behavior is the design working, not the design failing.

The demo video therefore uses explicitly labeled replay evidence — the UI badges its evidence mode at all times, per the design from day one. Degradation logs from the live pipeline are included in the repository.

Accomplishments that we're proud of

The system kept every quality gate on the critical path even when the model layer was failing. When live research became impossible on submission day, it refused to generate a market map from insufficient evidence — it logged an audit event and returned control to the user instead. Degrading honestly is harder to build than degrading silently.

What we learned

Multi-step LLM systems fail differently from normal software. Backend latency is usually predictable with low variance. Here, the same prompt against the same model varied by an order of magnitude within hours, and prompt weight affects latency super-linearly — the model decides how many search rounds to run, and you don't control that.

Timeout should trigger degradation, not retry. Retrying a slow call only makes it slower. This is obvious in hindsight and cost us a full request budget to learn.

Governance must stay on the critical path. Under time pressure, the tempting fix is to move quality gates off the hot path to shave seconds. We kept the Evidence Validator and constraint filtering synchronous throughout. The one component we allowed to go async was the alignment auditor — and we ultimately chose not to move even that.

Splitting one heavy call into two light ones beats tuning timeouts. A short grounded search plus a separate structured-extraction pass is far more tractable than asking one call to search, evaluate, and structure five fields for five products.

What's next for MarketCompass

  • Calibrate gate thresholds and the axis-selection correlation cutoff against real cases
  • Async alignment audit with post-hoc correction, now that the failure modes are understood
  • Multi-session memory so a returning user resumes an exploration rather than restarting it

Built With

Share this project:

Updates