Inspiration

Video calls lost some personality when Snap Camera went away. But the bigger problem is that calls still make small signals awkward. Saying "we're off track" interrupts the speaker. Planning-poker cards are hard to share over video. In class, I have waited for a pause to ask a question that never came.

I wanted those signals to live where people are already looking: on my video. Sceneuron makes a call easier to follow and a little more fun. Professional doesn't have to mean sterile.

What it does

Sceneuron puts live signs on your video. It works anywhere OBS's virtual camera works: Zoom, Meet, Teams, or a stream.

There are four ways to raise a sign:

  • Press a key on a Stream Deck. The Stream Deck is a physical keypad, and nothing could make it talk to Sceneuron, so I wrote a custom plugin for it in TypeScript. Each key looks like the card it raises and turns amber while that card is up.
  • Use your hands. Hold up three fingers and a "3" appears. Raise your hand and "Hand raised" goes up. This runs on the Mac itself, so it's free and nothing leaves your computer.
  • Say it. "Sceneuron, wrap it up." "My vote is eight." "Switch to the hacker theme." If it doesn't recognize a command, Gemini figures out what you meant.
  • Let the AI do it. Hold something up to the camera and a few seconds later it shows "Now showing: phone case."

It comes with four packs of cards: meeting signals (Wrap it up, We're off track, I have a question…), planning poker (0 to 21, ?, Coffee break), classroom (Hand raised, Please repeat, quiz answers A–D) and status (On mute, Back in 5, Presenting…). It also has a "Be right back" card that goes up by itself when you walk away from the camera.

The part I like most is the sprint-planning room. Everyone votes, and all you see is "✓ Voted · 3 of 5". Nobody can see a number until someone hits reveal. Then every card flips at the same moment, and you see everyone's vote and the average. If you're alone, a practice meeting adds fake teammates so you can try it.

It also narrates what it sees. A small bar on your video describes the scene as it happens: "1 person; right hand raised, peace sign; seen: coffee mug". That runs on your Mac on every frame, for free, and nothing gets sent anywhere. When you turn on scene understanding, Gemini's latest caption shows up under it too ("A person holding a phone case toward the camera · 4 s ago").

You can tell it what to watch for, in plain English. "I'm holding something up." "A pet walked in." When one of those becomes true, it shows "Now showing: coffee mug" or "Cat on stream!" It also knows when you leave the camera and puts up "Be right back" by itself.

Some other things it does:

  • Set up with AI: type "I run sprint planning on Zoom" and it builds the whole setup for you. Nothing changes until you approve it.
  • Live subtitles: your words show up under your video as you talk, translated into up to three languages at once. Since this is a hackathon, Klingon works too, written in real Klingon letters (pIqaD), plus Elvish in Tengwar and Toki Pona in sitelen pona. On a stream they can go out as real closed captions, and you can translate with a local model if you'd rather keep everything on your machine.
  • Reactions: ten of them, loud on purpose: "This meeting could have been an email", Mind blown, Ship it, Coffee break, Tongue out, You're absolutely right, We need more tokens, Works on my GPU, Meatproxy alert and My clanker will do it. Trigger them with a key, by voice, or with a facepalm on camera.
  • Themes: a clean default look, plus Windows 95, Hacker, 80s CRT, Classic Mac and Vaporwave, all one click away.
  • You stay in control: every change the AI makes shows up in a log with its reason and how much it cost, and an hourly budget caps the spend. A manual change always wins over the AI, and one key pauses everything.
  • Fits into what you already use: one click sets up OBS, the Stream Deck pairs by itself, and OBS hotkeys and Bitfocus Companion work too. You can save your whole setup as one file and share it.

How we built it

Sceneuron is an OBS plugin written in Rust. There were no maintained Rust bindings for OBS, so we wrote our own against the libobs C API. Camera analysis runs as a video filter on the webcam you already have, and OBS still renders a frame in about 6 ms with it on.

The AI is built in three layers, and the cheap ones go first. Calling a big model on every frame would be slow and expensive, so:

  1. Apple Vision runs on the Mac on every frame (around 21 ms). It finds faces, hands and their 21 joints, poses, objects and text. Hand signals, "Be right back" and the live description bar only use this layer, so they cost nothing.
  2. A quick judge, TypeSafe's Jev, only runs when the scene changes. It judges each condition from the text facts the vision layer found, in under a second, and never sees an image. When it's unsure, it asks for a full look.
  3. Gemini only looks at a frame when the cheap layers can't decide, or when something needs a name. Its answers are structured JSON that we validate. It also plans setups from one sentence, and Gemini 3.5 Flash-Lite translates the subtitles. Every call is priced, there's an hourly spending cap, and the cap is shown live.

ElevenLabs:

  • Scribe v2 Realtime turns your voice into commands and subtitles about 150 ms behind you.
  • We pass the card names in as keyterms, so it hears "off track" correctly.
  • It only acts on a clear match. "I don't agree" doesn't raise "Agree".
  • Flash v2.5 lets cards speak out loud. Each line is generated once and cached.

SpacetimeDB runs the poker room.

  • The module keeps rooms, participants and private ballots.
  • Votes sit in a table nobody can read until the reveal.
  • The reveal goes out to everyone at once, so the flip is synchronized.
  • We wrote our own client against SpacetimeDB's WebSocket protocol, because the official SDK's license doesn't mix with OBS's GPL.

Everything else:

  • The Stream Deck plugin is TypeScript, with a ready-made three-page profile and automatic pairing.
  • The editor is a local web app that also works as an OBS dock. It's usable by keyboard and screen reader.
  • I designed the look in Figma over eight versions and planned the build in Notability.
  • I planned the build with Claude Fable 5.1 on UltraCode, then implemented it with GPT-6 Astra on ultra effort with ultrafast on. That's how a project this size came together this fast.

Challenges we ran into

  • Making the AI cheap enough to leave on. That's why the layers exist. The first version asked the model way too often.
  • Five things changing the same overlay at once: the camera, voice, AI, Stream Deck and editor. The rule we settled on is that a person always beats automation. A hidden vote also must never flash on screen before the reveal, even when someone raises a card at the exact moment someone else reveals.
  • Testing at 3 AM. I held up a Powerade and nothing happened, because the quick check never saw the bottle and never asked Gemini. Then Gemini called it Gatorade because my hand covered the label. Now it names what's actually there: hold up an iced coffee and the sign says "iced coffee", not "plastic cup".
  • Video calls are blurry and mirrored. The cards had to be readable after Zoom compresses them, and the text had to read correctly for everyone else, not just in my own mirrored preview.

Accomplishments that we're proud of

  • It runs live in a real Zoom call today.
  • It passes 79 of 79 end-to-end checks inside a real OBS. There are over 1,900 automated tests, and the shared room passes all 17 steps of its live test against a real SpacetimeDB server.
  • Subtitles translate into three languages at once while you're still talking.
  • The design took eight rounds to get right. We ended up with calm, readable signs and reactions that are loud on purpose.

What we learned

  • Use the cheapest intelligence first. The big model is the last resort, not the default.
  • People remember the fun stuff (reactions, themes), but the useful stuff is why they keep it on. You need both.
  • Design for how your product is actually seen: small, blurry and mirrored on someone else's screen, not how it looks in a mockup.

What's next for Sceneuron

  • Windows and Linux support.
  • Poker rooms hosted on SpacetimeDB's cloud, so a whole team can join from anywhere.
  • Joining a room from your phone.
  • More packs: standups, retros, classrooms and streamers.

Built With

Share this project:

Updates

Submission history