Inspiration

Many people experience emotional stress quietly and informally—often without the energy to fill out forms, track moods, or seek immediate professional help. Most existing mental health tools feel either too clinical, high-effort, or reactive, addressing problems only after they escalate.

We were inspired to build a system that supports everyday emotional check-ins—something that listens first, reflects gently, and notices patterns over time without attempting to diagnose or replace professional care.


What it does

Talk With Zeno is a voice-first AI companion that allows users to express emotions naturally through speech or text.

The system:

  • Accepts quick, low-friction voice or text inputs
  • Tracks emotional and behavioral patterns across sessions
  • Responds with calm reflections, prompts, or journaling-style responses
  • Encourages healthier emotional awareness without clinical language

Zeno does not diagnose, score mental health, or provide medical advice. It is designed as a reflective companion, not a therapist.


How we built it

The system follows a clear frontend–backend architecture with explicit ownership of memory and safety.

Frontend (React + TypeScript)

  • Captures voice or text input
  • Streams audio chunks to the backend
  • Displays live transcription and synthesized responses

Backend (Python + Flask)

  • Orchestrates Speech-to-Text, LLM, and Text-to-Speech services
  • Manages explicit user memory and pattern tracking
  • Enforces safety and ethical constraints

AI Services

  • Speech-to-Text via Google Cloud Speech-to-Text
  • Language responses via Gemini (with fallback models)
  • Text-to-Speech via Groq Orpheus TTS (primary) with Google Cloud Text-to-Speech as fallback

The system was developed using Gemini 3 through the Antigravity IDE, which enabled rapid iteration, architectural experimentation, and tight integration between system logic and AI workflows.

A key design choice was keeping the language model stateless while handling memory and personalization explicitly in the backend.


Challenges we ran into

The primary challenge was the lack of a true streaming API for Speech-to-Text that could cleanly handle continuous voice input.

Because of this limitation, we had to batch audio data from the frontend to the backend, which introduced several issues:

  • API request–response cycles interrupting incoming audio data
  • Server overload risks during rapid chunk submissions
  • Improper chunk fragmentation and merge failures
  • Cancellation of in-flight STT requests when new data arrived

Designing a stable workaround required careful buffering, chunk sizing, retry logic, and explicit thresholds to ensure transcription remained usable without overwhelming the backend.

Another challenge was implementing a custom animated background effect that enhanced the calming tone of the product without distracting from the interaction. The butterfly animation was designed to follow a sinusoidal motion path with varying size and depth, requiring careful tuning of motion parameters to keep the effect smooth, subtle, and performance-friendly across devices.


Accomplishments that we’re proud of

  • A fully working streaming voice → STT → LLM → TTS pipeline
  • Explicit, backend-controlled longitudinal memory and personalization
  • Real-time transcription with graceful degradation
  • Clear ethical boundaries enforced outside the model
  • A usable, end-to-end prototype built under hackathon constraints
  • A custom, non-intrusive animated UI element that reinforces a calm, reflective user experience

Most importantly, we built a system that feels calm, intentional, and non-clinical, despite its technical complexity.


What we learned

  • Streaming voice systems are far more complex than batch-based pipelines
  • Managing memory outside the LLM provides better control and safety
  • Latency, chunking strategy, and request orchestration matter as much as model choice
  • Ethical constraints must be enforced at the system level, not left to prompts alone
  • Small UI details, when thoughtfully engineered, can significantly affect user comfort

We also learned the importance of designing for failure modes, not just ideal flows.


What’s next for Talk With Zeno

Next, we plan to:

  • Refine and stabilize existing features
  • Tighten backend security and access control
  • Research and evaluate the most suitable LLMs for this use case
  • Fine-tune models specifically for reflective, non-clinical responses

Planned feature additions include:

  • Auto-journaling: converting conversations into structured journal entries
  • Personal reflection summaries over time
  • Mental health pattern dashboards (non-diagnostic)
  • Curated learning resources linked to observed patterns

The goal is to evolve Talk With Zeno into a responsible, supportive system that complements—not replaces—human care.

Built With

Share this project:

Updates

Submission history