Inspiration

I've always believed that intelligence, talent, and hard work should open doors. Throughout my own life, though, I realized that this isn't always how the world works.

I've felt that I missed opportunities because of my language skills, because I couldn't put my ideas into words, or simply because I didn't know how to stand out when communication mattered. And I kept asking myself: how many talented, intelligent people are experiencing the same thing?

School teaches us mathematics, science, and engineering. It rarely teaches us how to speak with confidence, handle a difficult conversation, sell an idea, or make a lasting impression. So the person who gets the opportunity isn't always the most knowledgeable one. It's the one who knows how to communicate their value.

I don't believe people should have to become naturally extroverted to succeed. Communication is a skill, and like any skill, it can be learned, practiced, and improved. That's why I created Atrium.

What it does

I didn't want another chatbot that gives advice. I wanted a place where people actually live the situation, make mistakes, get feedback, and try again.

Atrium is a 3D city where every building is a situation you can practice in. Inside, you don't type into a text box: you speak out loud, in real time, to people who answer with their own voice and a lip-synced face.

  • The Company. You upload your CV and the job description. AI agents research the company, its website, and interview insights online, then build the likely interview process round by round. You then take the interview:

    • one-on-one or with a panel;
    • online as a video call, or in person, from the front door and the reception to the meeting room. The Company also has rooms for sales pitches and investor meetings, negotiations, team meetings, leading a team call, presentations, and networking.
  • The University. Academic presentations, a thesis defence in front of a committee, oral exams, debates, and scholarship or admission interviews prepared from the program's requirements and your own application.

  • Practice Center. Short exercises that turn feedback into progress: communication frameworks, articulation, filler words, structure. Personal bests, challenges, and a daily streak keep you consistent.

  • Everyday life: the Park, the Café, the Bus Stop. Small talk, introducing yourself, saying no, disagreeing politely, emotional conversations with a friend, and a research-based Storytelling course.

  • Brain Lab. An AI neuroscientist that discusses stress, focus, or confidence, grounded in published research with named sources, ending with a personal plan.

  • Language Center. Live conversations in the language you're learning, with support in your native language, so you speak from day one.

After every session, Atrium gives detailed, actionable feedback. It points to the exact moments where an answer, its structure, or its delivery could be stronger, including body language when the camera is on. One tap on a suggestion opens the matching exercise in the Practice Center.

How I built it

Atrium is a system of specialized agents, each with one clear job: brief.

  1. Research and preparation agents read the user's documents, search the web, and build each situation: the interview rounds, the committee, the meeting brief.
  2. Live conversation agents hold the real-time voice conversation. Each room has its own cast and its own rules: one person or a whole panel, who may speak, and when.
  3. Evaluation agents transcribe the session, score it against a competency catalog, and write the feedback, with timestamps that link back to the recording.

The stack:

  • Frontend: React, Three.js for the 3D city, and Capacitor for the Android app.
  • Backend: Python and FastAPI on Google Cloud Run, with Firestore for data and Cloud Tasks for the feedback pipeline.
  • AI: Gemini Live for real-time voice through Google's Agent Development Kit, and Gemini for research and evaluation.
  • Faces: a lip-sync server on a GPU virtual machine. It paints each speaker's mouth onto filmed footage of the room and streams the result over WebRTC, with a TURN relay for networks that block video calls.
  • Payments: RevenueCat subscriptions, with a free trial of live minutes.

Challenges I ran into

  • Cost. I wanted live voice, lip-synced faces, research, and detailed feedback without infrastructure that burns money while idle. For any session I watch the cost per minute of conversation:

$$ c_{\text{min}} = c_{\text{voice}} + \frac{c_{\text{GPU}}}{60,n} + \frac{c_{\text{feedback}}}{T} $$

Here:

  • (c_{\text{voice}}) is the real-time voice cost per minute;
  • (c_{\text{GPU}}) is the GPU price per hour, shared by (n) simultaneous conversations;
  • (c_{\text{feedback}}) is the one-off cost of analyzing a session of (T) minutes.

An L4 GPU delivers approximately (f_{\text{rendered}} \approx 29) FPS, exceeding the required (f_{\text{needed}} = 25) FPS.

That leaves room for exactly one smooth conversation per GPU. This constraint shaped the entire scheduling design: one conversation per machine, with the user waiting on a loading screen rather than encountering a frozen face.

  • Natural turn-taking

A voice AI determines that you've finished speaking after a silence of duration (\tau). If (\tau) is too short, it interrupts you mid-thought. If it's too long, every reply feels like a hang.

I tuned (\tau) for each room, ranging from (1.2) seconds for quick small talk to (2.2) seconds for a formal debate. I also added additional safeguards on top:

  • No double questions: a second question can't play before the first one is answered.
  • No talking over the user: a reply never starts while the user is still speaking.
  • No surprise endings: a goodbye only ends the conversation when the time the user chose is almost up, or when the user is the one leaving.

  • One product, not ten tools. Each environment has a different purpose. Making them feel like parts of the same city meant one shared live-conversation engine, one feedback format, and one way into practice.

What I learned

  • A good AI product is a workflow, not a model. The quality came from giving each agent a clear responsibility and connecting them well, not from picking the biggest model.
  • Real-time is a different world. Small timing details, like when to answer, when to stay silent, and how far the face lags behind the voice, matter more to users than raw intelligence.
  • Measure, then decide. Logging reply latency, frame rates, and cost per minute turned arguments into numbers, and numbers into decisions.
  • Experiment and accept trade-offs. Some technically impressive approaches made the experience worse. Testing, rethinking, and simplifying were part of the work.

What's next for Atrium

Atrium started with a personal frustration, but the ambition is much bigger. I want to help:

  • people who have brilliant ideas but struggle to express them;
  • people who are highly knowledgeable but uncomfortable speaking in public;
  • people learning a new language;
  • people who simply want more confidence in everyday conversations.
  • people preparing for a big moment in a language that isn't their first.

Today, every room except the Language Center runs in English. Next, I want users to practice in their own language and in the languages of the jobs and schools they're aiming for: a job interview in French, a pitch in Spanish, a thesis defence in Korean. The people you talk to, the feedback, and the exercises will all follow the language you choose, because the opportunities people miss aren't only in English.

I want to help people stop missing opportunities simply because they haven't yet mastered the art of communicating their potential. Being intelligent should never be a disadvantage just because you're not the loudest person in the room.

Atrium is where you practice before the moment matters.

Built With

Share this project:

Updates

Submission history