Inspiration

Millions of talented professionals outside native English-speaking countries possess exceptional technical and analytical skills, yet hit a ceiling when interviewing with global companies. The barrier is rarely a lack of domain expertise—it is the challenge of articulating complex ideas in high-stakes English interviews under pressure, eliminating hesitant pauses, and adopting the exact vocabulary that international hiring managers expect.

We built TechVoz to bridge this gap: an AI-powered conversational interview coach that transforms how non-native speakers practice, polish, and master their professional English communication.


What it does

TechVoz provides an interactive, full-cycle mock interview simulation tailored to any role, seniority level, and job description:

  • Adaptive Question Engine: Tailors interview questions directly to custom job postings, technical stacks, or non-technical business analyst domains (e.g., discovery workshops, KPI dashboards, stakeholder conflict resolution).
  • Continuous Voice Interaction: Candidates speak directly into their microphone with persistent speech recognition that preserves transcription flow even during natural thinking pauses.
  • Multi-Tier Model Answers: Compares candidate answers against Junior, Mid-Level, and Staff/Principal benchmarks to illustrate the exact difference in phrasing, technical depth, and trade-off analysis.
  • Executive Polish & Vocabulary Upgrades: Instantly transforms conversational thoughts into crisp, executive-ready Silicon Valley pitches and highlights industry-standard terminology.
  • Spanish-to-English Fluency Bridge: Detects Spanish idioms or mental translations, providing bilingual explanations and targeted coaching drills to build real confidence.
  • Interactive Audio Playback: Realistic synthesized interviewer voices with adjustable playback speeds (0.9x to 1.2x) to sharpen listening comprehension.

How we built it

  • Frontend Architecture: React 18 with TypeScript, Tailwind CSS, and Lucide icons, organized into modular workspaces (setup wizard, active speech recorder, interactive feedback drawers, and vocabulary drills).
  • Speech Engine: Custom Web Speech API wrapper engineered with persistent buffer management to ensure zero text loss across speech recognition silence boundaries and audio pauses.
  • Backend & Intelligence: Express server powered by Google Cloud and the Gemini 3.7 Flash API (@google/genai), leveraging structured JSON response schemas for deterministic question synthesis, real-time speech evaluation, and multi-tier phrasing generation.
  • Resilient Fallback Layer: A contextual offline evaluation and question generator that ensures zero latency interruptions and strict alignment between the active question and executive coaching feedback.

Challenges we ran into

  1. Speech Recognition Silence Timeouts: Standard browser speech recognition often resets its buffer after a few seconds of silence, causing candidates to lose their transcribed thoughts mid-answer. We resolved this by building an incremental accumulation buffer that continuously concatenates historical segments.
  2. Context-Aware Seniority Calibration: Balancing questions across diverse profiles—from junior business analysts working with Excel/Power BI to Staff distributed systems engineers—required careful prompt engineering and validation logic so candidates receive strictly relevant scenarios.
  3. Low-Latency Feedback Synchronization: Coordinating real-time transcription, evaluation criteria, and multi-tier model answers while guaranteeing that feedback never drifts from the active topic under transient network delays.

Accomplishments that we're proud of

  • Delivering a fluid, end-to-end voice interview experience where candidates can practice without feeling rushed or judged.
  • Creating the "How a Staff Engineer would phrase it" and Executive Pitch coaching modules that offer immediate, actionable communication upgrades.
  • Designing a responsive, bilingual bridge that meets candidates where they are and accelerates their transition into fluent English communication.
  • Achieving a zero-flicker, resilient architecture that runs smoothly across both web and mobile viewports.

What we learned

  • The most impactful feedback for non-native speakers is not just correcting grammar, but providing high-leverage industry idiom upgrades (e.g., replacing "make it faster" with "minimize query execution cost and eliminate sequential scans").
  • Preserving conversational flow and handling pauses naturally is essential for reducing candidate anxiety during mock interviews.
  • Grounding AI evaluation in structured schemas ensures consistent, reliable scoring across technical depth, fluency, pacing, and confidence.

What's next for TechVoz

  • Live Multimodal Audio (Gemini Live API): Introducing real-time spoken interruptions, follow-up probe questions, and natural conversational back-and-forth directly over audio streams.
  • Company-Specific Interview Personas: Pre-configured interview tracks tailored to specific hiring cultures (e.g., Amazon Leadership Principles, Google Googlyness/System Design, McKinsey Problem Solving).
  • Personalized Fluency Analytics & Progress Tracking: Long-term metric dashboards tracking WPM pacing, filler word reduction, and vocabulary expansion across multiple practice rounds.
  • Team & BootCamp Collaboration: Enabling university career centers and tech bootcamps to assign tailored interview tracks and track student readiness metrics.

Built With

Share this project:

Updates

Submission history