Inspiration
Millions of talented professionals outside native English-speaking countries possess exceptional technical and analytical skills, yet hit a ceiling when interviewing with global companies. The barrier is rarely a lack of domain expertise—it is the challenge of articulating complex ideas in high-stakes English interviews under pressure, eliminating hesitant pauses, and adopting the exact vocabulary that international hiring managers expect.
We built TechVoz to bridge this gap: an AI-powered conversational interview coach that transforms how non-native speakers practice, polish, and master their professional English communication.
What it does
TechVoz provides an interactive, full-cycle mock interview simulation tailored to any role, seniority level, and job description:
- Adaptive Question Engine: Tailors interview questions directly to custom job postings, technical stacks, or non-technical business analyst domains (e.g., discovery workshops, KPI dashboards, stakeholder conflict resolution).
- Continuous Voice Interaction: Candidates speak directly into their microphone with persistent speech recognition that preserves transcription flow even during natural thinking pauses.
- Multi-Tier Model Answers: Compares candidate answers against Junior, Mid-Level, and Staff/Principal benchmarks to illustrate the exact difference in phrasing, technical depth, and trade-off analysis.
- Executive Polish & Vocabulary Upgrades: Instantly transforms conversational thoughts into crisp, executive-ready Silicon Valley pitches and highlights industry-standard terminology.
- Spanish-to-English Fluency Bridge: Detects Spanish idioms or mental translations, providing bilingual explanations and targeted coaching drills to build real confidence.
- Interactive Audio Playback: Realistic synthesized interviewer voices with adjustable playback speeds (0.9x to 1.2x) to sharpen listening comprehension.
How we built it
- Frontend Architecture: React 18 with TypeScript, Tailwind CSS, and Lucide icons, organized into modular workspaces (setup wizard, active speech recorder, interactive feedback drawers, and vocabulary drills).
- Speech Engine: Custom Web Speech API wrapper engineered with persistent buffer management to ensure zero text loss across speech recognition silence boundaries and audio pauses.
- Backend & Intelligence: Express server powered by Google Cloud and the Gemini 3.7 Flash API (
@google/genai), leveraging structured JSON response schemas for deterministic question synthesis, real-time speech evaluation, and multi-tier phrasing generation. - Resilient Fallback Layer: A contextual offline evaluation and question generator that ensures zero latency interruptions and strict alignment between the active question and executive coaching feedback.
Challenges we ran into
- Speech Recognition Silence Timeouts: Standard browser speech recognition often resets its buffer after a few seconds of silence, causing candidates to lose their transcribed thoughts mid-answer. We resolved this by building an incremental accumulation buffer that continuously concatenates historical segments.
- Context-Aware Seniority Calibration: Balancing questions across diverse profiles—from junior business analysts working with Excel/Power BI to Staff distributed systems engineers—required careful prompt engineering and validation logic so candidates receive strictly relevant scenarios.
- Low-Latency Feedback Synchronization: Coordinating real-time transcription, evaluation criteria, and multi-tier model answers while guaranteeing that feedback never drifts from the active topic under transient network delays.
Accomplishments that we're proud of
- Delivering a fluid, end-to-end voice interview experience where candidates can practice without feeling rushed or judged.
- Creating the "How a Staff Engineer would phrase it" and Executive Pitch coaching modules that offer immediate, actionable communication upgrades.
- Designing a responsive, bilingual bridge that meets candidates where they are and accelerates their transition into fluent English communication.
- Achieving a zero-flicker, resilient architecture that runs smoothly across both web and mobile viewports.
What we learned
- The most impactful feedback for non-native speakers is not just correcting grammar, but providing high-leverage industry idiom upgrades (e.g., replacing "make it faster" with "minimize query execution cost and eliminate sequential scans").
- Preserving conversational flow and handling pauses naturally is essential for reducing candidate anxiety during mock interviews.
- Grounding AI evaluation in structured schemas ensures consistent, reliable scoring across technical depth, fluency, pacing, and confidence.
What's next for TechVoz
- Live Multimodal Audio (Gemini Live API): Introducing real-time spoken interruptions, follow-up probe questions, and natural conversational back-and-forth directly over audio streams.
- Company-Specific Interview Personas: Pre-configured interview tracks tailored to specific hiring cultures (e.g., Amazon Leadership Principles, Google Googlyness/System Design, McKinsey Problem Solving).
- Personalized Fluency Analytics & Progress Tracking: Long-term metric dashboards tracking WPM pacing, filler word reduction, and vocabulary expansion across multiple practice rounds.
- Team & BootCamp Collaboration: Enabling university career centers and tech bootcamps to assign tailored interview tracks and track student readiness metrics.
Built With
- ai
- career-coaching
- cloud-run
- edtech
- english-learning
- express.js
- full-stack
- gemini-3.7-flash
- generative-ai
- google-cloud
- google-gemini
- llm
- lucide-react
- mock-interview
- natural-language-processing
- node.js
- react
- rest-api
- speech-recognition
- tailwindcss
- typescript
- vite
- voice-ai
- web-speech-api
Log in or sign up for Devpost to join the conversation.