Inspiration

When I brainstormed project ideas with ordinary AI chatbots, the feedback often felt too agreeable. It would make an idea sound impressive without asking the uncomfortable questions: Does anyone actually need this? Can it be built this weekend? Is the “nobody else does this” claim even true?

I wanted the opposite experience: a pressure test that feels like pitching to real people. That became Genie Jury—four cloud-seated AI jurors with different personalities, priorities, and voices. They can interrupt a pitch, challenge assumptions, research risky claims in a live browser, disagree with each other, and still leave the builder with a constructive recovery plan.

What it does

A builder pitches an idea by voice or typed fallback. Genie Jury brings in four specialists:

  • Ember, the Builder challenges feasibility, scope, and what can actually ship.
  • Gale, the Skeptic checks factual and competitive claims using a real Browserbase cloud browser.
  • Tide, the User focuses on who has the problem, how they solve it now, and why they would switch.
  • Volt, the Jester delivers a playful roast that exposes the clearest flaw without being cruel.

The jury does not just give generic opinions. A Bailiff agent breaks the pitch into claims and assigns non-overlapping investigation angles. Gale can gather real evidence, while the other jurors use the shared evidence and each other’s findings to form independent assessments. At the end, the Ally provides a practical “cut, prove, test, build” plan, alongside four scorecards out of 100.

How we built it

We built Genie Jury as a React and TypeScript experience with a full-screen daytime-sky stage, illustrated jurors, speaking states, live captions, voice cues, scorecards, and a Browserbase research dock.

The backend is a local Node API that keeps provider credentials off the client. OpenAI Realtime supports the live microphone experience, while OpenAI Responses powers claim extraction, agent reasoning, jury interruptions, and structured deliberation. Browserbase provides live web search, cloud browser sessions, screenshots, and Stagehand actions for evidence gathering. ElevenLabs gives each juror a distinct voice and personality.

The agents coordinate through a shared evidence ledger and an inspectable run tree, so the final verdict can be traced back to real claims, sources, tool calls, and agent handoffs.

Challenges we ran into

The hardest part was making the jury feel like a conversation instead of four chatbot responses in a row. We had to handle interruptions in both directions: jurors can cut into a pitch, but the builder can also talk over a juror and take the floor back.

We also had to make the system honest under failure. Browser research can time out, voice generation can fail, microphones can be denied, and a user might only say “Can you hear me?” instead of pitching an idea. Genie Jury now has explicit pitch validation, provider fallbacks, bounded timeouts, safe exits, typed fallback, and no fabricated evidence or fake verdicts.

Another challenge was visual: the jurors had to feel dominant and readable without covering the interface. We rebuilt the stage around a bright sky, cloud-seated character art, subtle active-speaker motion, and minimal controls so the pitch still feels like a real panel.

Accomplishments we’re proud of

  • Built a live, interruptible multi-agent pitch conversation instead of a static feedback form.
  • Made Browserbase research visible on stage, including sources, screenshots, and evidence status.
  • Kept evidence honest: claims are marked verified, contested, or unproven rather than invented.
  • Created four distinct agent roles, voices, scoring rubrics, and challenge styles.
  • Designed recovery behavior for invalid pitches, microphone denial, provider failures, and early session exits.
  • Turned harsh feedback into an actionable plan instead of leaving users discouraged.

What we learned

We learned that multi-agent systems are more convincing when each agent has a clear responsibility, limited tools, and a shared source of truth. More agents alone do not create better feedback; coordination, evidence boundaries, and timing do.

We also learned that “AI-powered” is not enough for a good demo. The important part is making the AI’s work visible: showing the browser research, agent handoffs, sources, and reasoning trail gives the audience something they can inspect and trust.

What's next for Genie Jury

Next, we want to add saved pitch sessions, richer evidence comparison, selectable jury personalities, stronger real-time turn-taking, and a post-verdict practice mode where builders can retry their pitch after applying the Ally’s feedback.

Longer term, Genie Jury could become a practice room for hackathon teams, student founders, startup accelerators, and anyone who needs honest feedback before pitching in a real room.

Built With

Share this project:

Updates

Submission history