Inspiration

We all have moments when we need guidance—before a job interview, learning a new language, or just processing a stressful day. But coaches, tutors, and mentors aren't available 24/7.

What if your phone contacts included AI experts you could actually call anytime? Not chatbots with text bubbles, but voices you recognize—assistants that feel like real people in your pocket.

And the big one: traveling abroad and wishing you could speak the local language in your own voice? That sparked the real-time translator with voice cloning.

What it does

Vox turns AI into voice contacts you can call anytime.

  • AI Contacts: Interview coach Alice, Spanish tutor Carlos, wellness mentor Dr. Sam—each with unique voices and personalities
  • Push-to-Talk: Hold to speak, release to send. Natural voice conversation flow
  • Voice Cloning: Record 30 seconds of your voice, clone it, and use it across contacts
  • Real-Time Translator: Speak in English, hear yourself respond in Spanish, French, German, Japanese—30+ languages in YOUR voice
  • Multi-AI Support: Choose between Gemini, Claude, GPT-4, or DeepSeek per contact
  • Works Everywhere: Web, PWA, and Telegram Mini App

How we built it

Layer Tech
Frontend Next.js 16, React 19, TypeScript, Tailwind CSS
AI Brain Google Cloud Vertex AI (Gemini 2.0), Claude, GPT-4, DeepSeek
Voice ElevenLabs TTS, STT, Voice Cloning API
Backend Firebase Auth, Firestore, Cloud Storage
Observability Datadog LLM Observability, RUM, Detection Rules
Infrastructure Upstash Redis (rate limiting), Stripe (payments)

The architecture prioritizes voice-first UX—minimal latency from speech to AI response to spoken reply.

Challenges we ran into

  • Voice latency: Getting under 2 seconds from user speech → AI response → TTS playback required aggressive optimization and streaming
  • Conversation context: Maintaining natural dialogue across voice turns without losing context
  • Voice cloning quality: Ensuring 30 seconds of audio produces a recognizable, natural-sounding clone
  • Multi-provider routing: Seamlessly switching between Gemini, Claude, and GPT-4 while maintaining consistent UX
  • Telegram Mini App constraints: Adapting the full experience to work within Telegram's webview limitations

Accomplishments that we're proud of

  • Voice cloning that actually works: Users hear themselves speaking languages they don't know
  • Sub-3-second response loop: Speak → AI thinks → Voice responds, feels like a real call
  • Enterprise-grade observability: Full Datadog LLM monitoring with detection rules that auto-create incidents
  • 5 ready-to-use AI contacts: Each with distinct personality, voice, and purpose
  • Cross-platform: Same experience on web, mobile PWA, and Telegram

What we learned

  • Voice UX is fundamentally different from chat UX—timing and tone matter as much as content
  • LLM observability isn't optional at scale—you need to track latency, costs, and errors from day one
  • ElevenLabs voice cloning is surprisingly good with just 30 seconds of clean audio
  • Users form emotional connections with AI voices faster than with text-based assistants

What's next for Vox

  • Group Calls: Mix AI contacts with real friends—imagine a study group with your friend, an AI tutor, and GPT-4 brainstorming together
  • AI Music Creation: Contacts that compose and send personalized music to friends
  • Memory & Relationships: Each contact remembers your conversations, preferences, and builds a relationship over time—like a real friend who knows your history
  • Voice Avatars: Video avatars that lip-sync with the AI voice for video calls
  • Enterprise Version: Custom AI contacts for company onboarding, training, and support

Pitch Deck: https://vox-aicontact-fe0e3.web.app/pitch-deck

Built With

  • anthropic-claude
  • app
  • datadog
  • deepseek
  • elevenlabs
  • firebase-auth
  • firebase-storage
  • firestore
  • framer-motion
  • gemini-2.0
  • google-cloud-vertex-ai
  • gsap
  • mini
  • next.js
  • openai
  • pino
  • react
  • stripe
  • tailwind-css
  • telegram
  • typescript
  • upstash-redis
  • zustand
Share this project:

Updates

Submission history