Inspiration
Most mental-health apps stop at text chat. But when someone is really struggling, a screen feels cold — what helps is a voice. We wanted to know: can an AI companion cross that line, from chat to an actual phone call, without becoming a fake therapist? That question shaped everything else.
What it does
- CBT-informed chat with two personas: Maya (empathetic, warm) and Liam (analytical, grounding). Pick who you want to talk to; both stream replies and keep talking even if the model backend drops (offline fallback). Voice input + read-aloud via iFlytek, browser fallback.
- CBT reframing, two depths: Quick CBT Studio takes a thought you paste and names the distortion behind it — all-or-nothing, catastrophizing, mind reading, emotional reasoning — then returns a balanced reframe. The Guided Journey walks you through the same process step by step, assess first, reframe after, for when one-shot feels too fast.
- CALL-E phone companion: tap "Call Me", enter a number (E.164), consent, and the AI rings you — Maya calls for a live 5–10 minute voice check-in with validation and grounding exercises. You can minimize the dialing modal and keep typing while she calls. Post-call, structured results (outcome, mood change, support summary) come back to the UI. The number exists only in-flight — never stored.
- Somatic grounding: a kinetic mandala breath guide with 7 evidence-based patterns — 4-7-8 relaxation, Box 4-4-4-4 focus, Coherent 4-4 for HRV, physiological sigh for instant relief, energy breath, triangle zen, and 5-4-3-2-1 sensory grounding — paced by Web Audio singing-bowl cues, not a silent timer.
- Dual-axis mood mapping: plot how you feel on energy × valence, tag the emotion, leave a note, then one click carries that mood into the chat — the entry point to talking, not a dead-end chart.
- Dialogue audit:
/api/analyzereviews recent conversation for emotional climate, recurring patterns, and one concrete growth step — so reflection compounds instead of resetting each session. - Crisis safety gateway: high-risk phrasing is caught locally and server-side before any model call; the user instantly gets real hotlines (988 + country-specific) instead of a generated maybe-answer.
How we built it
React 19 + TypeScript + Vite + Tailwind v4 on the client, GSAP for the fluid motion. Backend is one Tencent CloudBase serverless Express function: /api/chat, /api/reframe, /api/analyze, and the CALL-E proxy at /api/call/*. All secrets live server-side, so the client has zero keys. Speech uses iFlytek with Web Speech fallback.
Challenges we ran into
- Timeout chain: CALL-E's upstream review takes ~15–20 s, so we had to size $t_{\text{client}} > t_{\text{proxy}} > t_{\text{upstream}}$ (50 s > 45 s) — a shorter timeout reported false failures on calls that actually went through.
- Safety vs. personality: making the crisis interceptor bypass inference entirely without breaking the conversational flow.
- Privacy by design: phone numbers exist only in-flight, never in a DB or log.
- Region gotchas (+86 rejected up front), per-IP quotas, and mapping upstream error codes to messages humans understand.
Accomplishments that we're proud ofconversational flow.
- Privacy by design: phone numbers exist only in-flight, never in a DB or log.
- Region gotchas (+86 rejected up front), per-IP quotas, and mapping upstream error codes to messages humans understand.
Accomplishments that we're proud of
The moment the demo phone actually rang — that first live CALL-E check-in is when MindQuark stopped being "another chatbot" and became something you can hear. Beyond that:
- A safety model we can defend: the crisis interceptor runs on both client and server and short-circuits inference entirely on high-risk input, so a hotline — never a language guess — is what a user in crisis sees.
- Privacy that held up under audit: zero phone-number persistence, zero client-side secrets, daily per-IP quotas — we could open any log mid-hackathon and find nothing to leak.
- It survives its own dependencies: primary/backup LLM failover, offline fallback for chat, and browser Web Speech when iFlytek doesn't answer. One provider going down doesn't take the sanctuary with it.
- Full EN/CHINESE parity, typed per-key, on every page including the breath mandala and call flow — both languages feel first-class, not translated.
- Shipped, not demoed: live on CloudBase with a working gateway health check — judges can call their own phone, no localhost required.
What we learned
- Latency budgets are a product feature, not an implementation detail.
- In wellbeing tech, what you refuse to do (store data, answer crisis turns with model guesses) matters more than what you ship.
- Voice and text are different emotional channels — the call feels real in a way chat doesn't.
What's next for MindQuark
- Proactive check-in calls on user-set schedules, not just on demand
- Follow-up call summaries in the mood journal (stored locally only)
- More regions/languages for CALL-E, and call-based grounding exercises synced with the breath mandala
- Eval harness for reframe quality, so safety claims stay testable
Built With
- ai-chatbot
- call-e-api
- cloud-functions
- express.js
- framer-motion
- gsap
- i18n
- iflytek-asr
- iflytek-tts
- llm
- node.js
- openrouter
- react
- rest-api
- serverless
- tailwind-css
- tencent-cloudbase
- typescript
- vite
- voice-ai
- web-app
- web-audio-api
- web-speech-api
Log in or sign up for Devpost to join the conversation.