A MULTIMODAL DEEP LEARNING FRAMEWORK FOR REAL-TIME HUMAN EMOTION RECOGNITION USING FACIAL, VOCAL, AND PHYSIOLOGICAL SIGNAL FUSION FOR MENTAL HEALTH MONITORING AND INTELLIGENT HUMAN–COMPUTER INTERACTION

Authors

  • Muhammad Essa Siddique Author
  • Mushtaq Ahmed Author
  • Muhammad Umar Akram Author
  • Hameed Hussain Author
  • Ashraf Zia Author
  • Humaira Batool Author

Keywords:

Multimodal emotion recognition; affective computing; deep learning; cross-modal attention; facial expression analysis; speech emotion recognition; physiological signal processing; mental health monitoring; human–computer interaction; real-time inference.

Abstract

Reliable emotion recognition is essential for affective computing, mental health monitoring, and intelligent human–computer interaction; however, unimodal approaches remain vulnerable to environmental noise, signal uncertainty, occlusion, and ambiguity in human emotional expression. This study proposes a real-time multimodal deep learning framework that integrates facial video, vocal signals, and physiological measurements, including electroencephalography, electrocardiography, and galvanic skin response. The proposed architecture employs modality-specific feature learning, where facial dynamics are extracted through convolutional and temporal convolutional layers, vocal characteristics are modelled using a CNN–BiLSTM network, and physiological patterns are captured using one-dimensional convolution and temporal transformer modules. A cross-modal multi-head attention mechanism adaptively integrates heterogeneous representations, while contextual conditioning improves temporal affective modelling. The framework was evaluated using seven experimental configurations, including three unimodal, three bimodal, and one trimodal architecture, through comparative performance analysis, component-level ablation, confusion matrix evaluation, ROC–AUC assessment, and inference latency measurement.

The proposed trimodal framework achieved an accuracy of 91.7%, outperforming facial-only (76.4%), vocal-only (71.8%), and physiological-only (68.3%) models. Bimodal fusion achieved accuracies of 84.2% for facial–vocal, 82.6% for facial–physiological, and 79.5% for vocal–physiological combinations. Replacing adaptive attention fusion with conventional feature concatenation reduced accuracy to 85.9%, demonstrating the contribution of cross-modal attention learning. Ablation experiments further confirmed the importance of physiological temporal modelling, vocal sequential representation, facial alignment, and contextual refinement. The final model achieved macro precision, recall, and F1-scores of 0.889, 0.873, and 0.881, respectively, with a macro-ROC–AUC of 0.953. The framework contains approximately 21.3 million trainable parameters and processes each analysis window within 47 ms. The findings demonstrate the effectiveness of attention-driven multimodal fusion for robust emotion recognition while emphasizing the importance of further validation under diverse real-world conditions, privacy-preserving learning strategies, and cross-subject evaluation.

Downloads

Published

2026-09-17