SkyRoom: the meeting app with a teammate who remembers

Inspiration

Software teams spend hours in meetings and then lose most of what was said. Someone asks "didn't we decide this last sprint?" and nobody can find it. Someone else says a fix shipped when the PR is still open. Afterwards, someone has to type up notes and copy action items into Jira by hand.

Today's AI meeting tools are bots that join Zoom, record the call and send a summary later. They can't answer a question during the meeting, they don't know your repo or your tickets, and they can't tell what was decided from what was only discussed.

So we built our own meeting app, with an AI teammate built into it.

What it does

SkyRoom is a video meeting app for software teams. Its agent, Polaris, sits in every meeting. It only acts when someone asks it to.

During the meeting

  • Live captions per speaker. Each person's audio is transcribed separately, so every line has the right name and a timestamp.
  • Ask by voice or chat. Say "Polaris, what did we decide about the billing retry?" or type @Polaris in chat. Polaris looks through past meetings, GitHub (issues, PRs, CI checks, code) and Jira, then answers with sources such as meeting 3 · 14:02, dropsubs/website#7 or DS-12.
  • No surprise speech. A voice question gets an answer card that everyone sees. Polaris speaks only when someone presses Speak. Anyone can also post the answer to chat or dismiss it.
  • Private questions stay private. A private question gets a private answer, and private chat is never saved, indexed or summarized.
  • Private fact checks. If someone says something the team's records contradict ("the fix is in 2.4", but the PR isn't merged), Polaris sends that person a private message with the sources. The rest of the room doesn't see it.
  • Agenda tracking. Polaris can draft the agenda from unfinished work and earlier meetings. It ticks items off as the conversation covers them.
  • Catch-up for late joiners. Polaris sends them a private summary of what they missed.
  • Live translation (optional). If the host turns it on, speech in other languages is captioned and saved in English.

After the meeting

  • A report with an overview, the decisions (each linked to the moment in the transcript where it was made), open questions and blockers.
  • Task drafts with owners and due dates only when someone actually said them. Polaris never guesses either.
  • Push to Jira after approval. An admin reviews the drafts and presses Push, and each one becomes a real Jira issue. Nothing is written to GitHub or Jira without a person approving it.
  • Every meeting goes into team memory, so the next meeting can ask about it.

How we built it

browser ── /       ──> board (Next.js)
        ── /api/*  ──> brain (FastAPI) ──> Postgres + pgvector
                                       ──> GitHub / GitLab / Jira MCP servers
                                       ──> Gemini (OpenRouter fallback), ElevenLabs
        ── WebRTC  ──> LiveKit Cloud <── realtime worker ──> brain /internal
  • board (Next.js, React, TypeScript, Tailwind, LiveKit Components): the meeting room, lobby, reports, tasks and settings.
  • realtime (Python LiveKit Agents worker): joins every room as Polaris. It runs one ElevenLabs Scribe v2 Realtime stream per participant, which is how captions get the right speaker without diarization. It sends final segments to the brain, detects the wake word ("Polaris" anywhere in a sentence, waiting if the question trails off), and plays ElevenLabs TTS into the room when someone presses Speak.
  • brain (FastAPI, async): handles accounts, meetings and the agent.
    • Orchestrator: a Gemini tool-calling loop over small typed, read-only tools (meeting memory search, GitHub, GitLab, Jira, code search). Every answer comes back as a typed Answer with its evidence attached, plus a list of any sources that failed, so it says "Jira was unreachable" instead of making up a status.
    • Meeting memory: transcripts are cut into windows of about 1,200 characters (roughly 90 seconds of talk), tagged with speaker and timestamp, embedded with Gemini and stored in pgvector. Polaris's own words are never stored as evidence, so it can't end up citing itself.
    • Post-meeting pipeline: a summary, structured extraction of decisions and tasks, and decision chains across meetings, all validated against Pydantic schemas.
    • Fact checker: a cheap filter first picks out sentences that look like claims, and only those go to the LLM, which compares them with GitHub and Jira.
    • Agenda tracker: a small, fast decision model on OpenRouter checks after each caption whether an agenda item is done.
  • Integrations through MCP: GitHub uses GitHub's hosted MCP server with a fine-grained read-only token, encrypted at rest. Jira pushes go through the REST API with the team's own account. For the demo we built mock GitHub, GitLab and Jira MCP servers over a fictional company called DropSubs, so judges can try it without our credentials.
  • Model fallbacks: if Gemini returns 429 or 5xx, the brain tries backup Gemini models and then OpenRouter. Embeddings fall back the same way.
  • Deployment: one Docker Compose stack (Postgres, migrations, brain, worker, mocks, board, Caddy) on one server.

Challenges we ran into

  • Knowing when Polaris is being asked. Running the LLM on every sentence is slow, expensive and annoying. We match the wake word on final transcripts instead, handle fillers like "Um, Polaris", and wait briefly when a question isn't finished.
  • Speaker identity. Mixed-audio diarization guesses who is talking. Giving each LiveKit track its own STT stream means we always know.
  • Answers that cite themselves. Early versions quoted Polaris's own earlier answers as proof. We now keep the agent's speech out of memory completely.
  • Free-tier limits. Gemini's free embedding quota shaped our chunk size and made us build the OpenRouter fallback.
  • Keeping public and private separate. Every invocation records where it came from (voice, public chat or private chat), and the answer goes back only to that channel. Tests enforce this.

Accomplishments we're proud of

  • A working meeting app built from scratch in a weekend, not a bot added to Zoom.
  • Answers that cite real sources, and that say so when a source is missing or down.
  • A human approves anything before it reaches Jira.
  • Live end-to-end tests: a synthetic participant speaks a recorded question into a real LiveKit room, and we check the transcript and the answer.

What we learned

  • With agents, the hard part is permissions (who can trigger it, who sees the answer, what it can write), not the model.
  • Typed tool results and schema-validated output made Gemini much more reliable than prompt tweaking.
  • MCP let us switch between mock and real GitHub by changing one URL.

What's next for SkyRoom

  • Small coding tasks: the host approves the scope, Polaris opens a PR, and a person reviews the diff.
  • Linear and Slack connectors.
  • Removing members and per-meeting access controls.
  • Self-hosted LiveKit for teams that can't use the cloud.

Built With

Share this project:

Updates

Submission history