💡 Inspiration

Understanding human emotion is the next frontier in natural Human-Computer Interaction (HCI). Standard AI systems process what we say, but often miss how we feel. We built Emotion AI to bridge this empathy gap—creating an intelligent, real-time multimodal system capable of understanding human emotional states through visual cues and vocal acoustics. From improving adaptive e-learning platforms to providing real-time mental well-being feedback, Emotion AI brings genuine empathy to machine intelligence.

⚙️ How We Built It

Emotion AI leverages a Multimodal Deep Learning Architecture that combines computer vision for facial expression analysis with audio processing for vocal sentiment:

  1. Visual Pipeline (Facial Expression Recognition):

    • Built on a deep Convolutional Neural Network (CNN) leveraging custom residual blocks (ResNet backbone) to extract spatial hierarchies from facial keypoints.
    • Input images are normalized and mapped through convolutional kernels calculated by: $$O = \left\lfloor\frac{I - K + 2P}{S}\right\rfloor + 1$$
    • Feature maps pass through GELU activation layers and Batch Normalization to prevent internal covariate shift and ensure stable convergence.
  2. Audio Pipeline (Speech Emotion Recognition):

    • Raw audio signals are processed into Mel-Frequency Cepstral Coefficients (MFCCs) and spectrograms using Librosa.
    • Sequential temporal dynamics are processed using a Gated Recurrent Unit (GRU) / LSTM architecture, capturing pitch, cadence, and vocal intensity across time.
  3. Late-Fusion Multimodal Classifier:

    • Predictions from both the visual stream ($P_{\text{visual}}$) and acoustic stream ($P_{\text{audio}}$) are combined using a weighted late-fusion mechanism: $$P(\text{emotion}) = \text{Softmax}(\alpha \cdot z_{\text{visual}} + (1 - \alpha) \cdot z_{\text{audio}})$$
    • Outputs a probability distribution across 7 primary emotions (Happy, Sad, Angry, Fear, Surprise, Disgust, Neutral).
  4. Frontend & Real-time Stream:

    • High-throughput WebSockets streaming webcam frames and audio segments to a FastAPI backend optimized with ONNX Runtime for sub-50ms inference latency.

🚧 Challenges We Faced

  • Real-time Latency & Synchronization: Processing video frames (30 FPS) simultaneously with streaming audio packets caused initial lag. We solved this by using asynchronous queueing in Python and exporting trained PyTorch models to quantized ONNX format.
  • Class Imbalance: Datasets like FER-2013 have severe class imbalances (fewer Disgust and Fear samples). We applied focal loss and automated data augmentation (rotation, noise injection, pitch shifting) to equalize performance.
  • Generalization Under Noisy Environments: Micro-expressions and variable lighting degraded single-frame classification. Implementing temporal smoothing over rolling frame windows drastically improved stability.

🎓 What We Learned

  • How multimodal late-fusion outpaces single-modality networks when dealing with masked expressions or muffled audio.
  • The practical trade-offs between depth, parameter efficiency, and real-time inference speed when deploying models to edge devices.
  • Designing human-centered interfaces that display AI confidence scores transparently without overwhelming the user.

🚀 What's Next for Emotion AI

  • Edge Computing Integration: Porting models to execute entirely client-side using TensorFlow.js / WebNN for maximum privacy.
  • Biometric Expansion: Integrating physiological inputs (e.g., heart rate variability from wearable devices) to enrich emotional contextual awareness.
  • Empathetic Conversational Agent: Connecting Emotion AI output directly to LLM prompts to dynamically adjust an AI assistant's tone and response style.

Built With

Share this project:

Updates

Submission history