Inspiration

As someone who deals with social anxiety, high-stakes communication like a critical job interview or a major project presentation can be incredibly stressful. My mind tends to race, I rush my words, and it is easy to lose composure. The problem with existing communication tools is that they only give you a transcript or feedback after the call is over, when it is too late to fix anything. I needed something to help me in the moment. I wanted to build a tool that acts like a real, supportive coach sitting in the room with you, nudging you in real-time to slow down, breathe, and project confidence.

What it does

Vocalis is a real-time AI communication coach. When you select a high-stakes scenario, it generates a tense, realistic workplace question. As you speak into the microphone, the app live-analyzes your audio stream and scores your telemetry on the fly: Pace, Energy (vocal vitality), and Delivery Presence (command and composure).

Instead of overwhelming you with a wall of text, Vocalis isolates the one thing holding you back and delivers a single, targeted nudge (e.g., "Slow down to stay calm and clear"). It then uses a multi-attempt loop to track your delta scores across retries, visually proving your real-time improvement.

How I built it

  • Backend: I built the server using Python and FastAPI, utilizing WebSockets to continuously stream and process audio data.
  • AI Agent: I used the Strands SDK to build the coaching logic and manage the state of the conversation.
  • Local Inference: To ensure absolute privacy and low latency, inference runs entirely locally using Ollama. I utilized the Qwen 2.5 7B model, which generates snappy, sub-3-second responses without relying on cloud APIs.
  • Frontend: I designed a premium, dark-mode SaaS interface using vanilla JavaScript, focusing heavily on a clean visual hierarchy for the live telemetry.

Challenges I ran into

The biggest hurdle was LLM latency. Initially, running a massive 23GB Qwen model locally took over two minutes per prompt, which completely killed the real-time experience. I absolutely refused to fake the metrics or use pre-curated responses for the hackathon. Instead, i solved it by implementing a dynamic progress clock UI to mask the generation time gracefully, and ultimately optimized our pipeline to run the much faster 7B model. I also had to untangle tricky WebSocket concurrency bugs where duplicate client connections were locking up the Strands agent.

Accomplishments that I'am proud of

I am incredibly proud of the visual design and the real-time integration. Bridging the gap between raw backend audio streaming and smooth frontend metric visualization was a massive technical win, especially getting those Pace and Energy tiles to react instantly to my voice. The app looks and feels like an enterprise-ready product that can genuinely help people manage their anxiety and build confidence.

What i learned

Managing state in a real-time WebSocket application is unforgiving. I also learned a tremendous amount about optimizing local open-weight models and how critical UI state management is when masking asynchronous AI generation.

What's next for Vocalis

I plan to integrate custom user-defined scenarios, add historical progress tracking to map growth over weeks rather than just immediate retries, and build out native local audio processing to detect filler words (like "um" and "uh") without relying on cloud APIs.

Share this project:

Updates

Submission history