-
-
MindMate’s emotion calendar helps users track patterns, changes, and possible emotional triggers over time.
-
MindMate detects multiple emotions, shows its confidence and key word contributions, and lets users correct the result.
-
MindMate analyzes custom date ranges to summarize emotional intensity and show how each detected emotion changes over time.
-
Explainable AI shows which words contributed to each detected emotion, making MindMate’s predictions transparent and understandable.
-
Users can email an emotional summary to a trusted professional, helping MindMate connect AI insights with human care.
-
MindMate’s peer network connects users through topic-based support forums for shared experiences and encouragement.
Inspiration
This research emerged from personal tragedy. My friend's mother passed away due to suicide from untreated depression. Watching her recovery and seeing how mental health struggles shatter families and communities transformed my understanding of why this work matters. Mental health disorders affect 1 in 4 adults globally, cases have risen 48.1% since 1990, and yet there are only 9 mental health workers per 100,000 people worldwide, dropping below 1 per 100,000 in low-income countries. Over half of people with mental health disorders never receive treatment due to diagnostic barriers, cost, and stigma. I wanted to build something that could help close that gap.
What it does
MindMate is an AI-powered mental wellness journaling platform that detects 16 concurrent emotions from unstructured journal text with over 98.9% accuracy, while explaining why it made each detection and how confident it is. Users write or speak a journal entry, and MindMate identifies emotional patterns (including mixed states like anxiety combined with excitement, which single-model systems typically miss), flags crisis language for self-harm, abuse, or violence and connects users to resources immediately, and visualizes emotional trends over time on a calendar view. Users can override any detected mood if they disagree. The platform also delivers evidence-based, tailored recommendations for the detected emotion, lets users send an analysis report directly to a therapist by email, and connects users to moderated peer support boards. It has processed 72,817 diary entries during development.
How we built it
Application layer. The frontend and application logic run on .NET MVC in C#, handling authentication, the journaling interface, the calendar view, and the peer support boards. Journal entries, emotion scores, and user profiles persist to a database queried by date range for the longitudinal visualization.
Dataset engineering. I built a training set of 72,817 diary entries, engineered for diversity and balance: a 0.486 overall diversity score (0.990 emotion balance, 0.477 structural variation), 6,395 unique vocabulary terms, an average of 33 words per entry, 20.6% of entries containing 2 or more concurrent emotions, and age stratification from 13 to 65+.
Feature engineering. I built 13 custom linguistic features on top of the raw text, including sentiment indicators, self-reference frequency, intensity keywords, punctuation patterns, and case sensitivity, to make the models context sensitive.
Consensus architecture. I trained 11 models spanning gradient boosting (XGBoost, LightGBM), deep learning, ensemble methods (stacking, voting, bagging), linear models (L2-regularized logistic regression), and traditional ML (random forest, naive Bayes, SVM, decision trees). The consensus layer combines the top 5 by F1-score: Neural Network, Ensemble-Stacking, Ensemble-Voting, LightGBM, and Logistic Regression. Each model's prediction for a given emotion is weighted by that model's own F1-score for that emotion:
weight[model_i] = F1_score[model_i][e] consensus_score[e] = Σ(prediction[model_i][e] × weight[model_i]) / Σ(weight[model_i]) confidence[e] = (models_predicting_e / total_models) × consensus_score[e]
This avoids a flat average that would cancel out a model catching a signal the others miss. The consensus model reached 98.93% accuracy, 0.9307 F1-score, and 0.0107 Hamming distance, outperforming every individual model, with the largest gains in multi-emotion detection specifically.
Explainability. I built two layers: a feature-level layer using TF-IDF weights to show which phrases drove a detection, and a consensus-level layer showing model agreement and weighted confidence.
Crisis detection. A separate pass screens each entry for self-harm, abuse, and violence patterns, independent of the emotion pipeline, and immediately surfaces crisis resources when triggered.
Voice input. Whisper handles voice-to-text transcription so entries can be spoken instead of typed, feeding into the same 11-model pipeline as text entries.
Fast inference. Running 11 models per entry through a freshly spawned Python process added several seconds of latency from model load time alone. I built a persistent Python server that loads all 11 models into memory once and stays running as a background service, with the .NET MVC app sending requests to it over local calls, cutting latency to near-instant.
Challenges we ran into
The hardest technical challenge was building the consensus weighting system. Giving one model too much weight caused the ensemble to inherit that model’s errors, while equal weighting ignored the fact that some models performed better on specific emotions. I evaluated all 11 models separately for each emotion using precision, recall, and F1 score. I then used the per-emotion F1 scores to weight the five strongest models and tuned the decision thresholds to reduce multi-label errors.
Inference speed was another issue because each journal entry had to pass through multiple models and the feature-engineering pipeline. Loading every model for each request created too much latency, so I moved them to a persistent server architecture where the models remain in memory between requests.
The safety-detection system also required careful threshold tuning. I prioritized recall so the system would recognize more concerning language, but an overly sensitive threshold produced too many false positives. I repeatedly tested the precision-recall tradeoff to improve detection while preventing unnecessary flags from weakening user trust.
Accomplishments that we're proud of
The consensus model reached 98.93% accuracy and a 0.9307 F1-score across 16 emotions, beating every individual constituent model, including the standalone neural network and ensemble-stacking models. It performed especially well on multi-emotion detection, where single models plateau, which was the core goal of the project. I also validated the approach with Washington State therapists, the UVA Cognition Lab, and the UC Irvine Psychology department, who confirmed the explainability features build trust in the model's decisions and that the 16-emotion taxonomy reflects real psychological frameworks.
What we learned
Ensembling multiple heterogeneous models, weighted by per-class F1-score, produces meaningfully more reliable results than any single model for a task as nuanced as concurrent emotion detection. I also learned that clinical credibility matters as much as technical accuracy. Getting feedback from therapists and cognition researchers shaped decisions I wouldn't have made from the ML side alone. And building anything that touches crisis detection carries real weight. It's not just a feature to tune for a metric, it has to be designed with genuine care for the person on the other end.
What's next for MindMate: AI-Powered Mental Wellness
I'm following a phased plan. Phase 1 focuses on user-consented feedback loops for continuous retraining on novel linguistic patterns, plus partnerships with schools, communities, and therapists for broader adoption. Phase 2 integrates PHQ-9 and GAD-7 clinical screening tools to correlate MindMate's detections with clinical standards, alongside therapist registration for care coordination. Phase 3 explores multimodal emotion tracking through social media integration to capture emotions users don't explicitly journal about. Phase 4 adds an anonymized data sharing option to enable population-level mental health research.
Built With
- .net
- ai
- c#
- ensemble
- learning
- machine
- mvc
- natural-language-processing
- python
- whisper

Log in or sign up for Devpost to join the conversation.