A hierarchical multifunctional cross fusion network for emotion recognition from facial expressions and peripheral physiological signals

Objective. Accurate emotion recognition is essential for enabling adaptive and empathetic human–computer interaction. However, emotion recognition based on a single modality may be affected by modality-specific noise, weak emotional cues, and individual differences. This study aimed to develop a practical non-invasive emotion recognition framework by integrating facial expressions and peripheral physiological signals (PPSs). Approach. We investigated multimodal emotion recognition using facial images and wearable PPSs. To effectively integrate heterogeneous modalities, we proposed a hierarchical multifunctional cross-fusion network (HMCFNet). The framework first performs modality-specific feature refinement to extract emotionally relevant representations from each modality. A joint cross-attention module is then employed to enhance complementary cross-modal information while suppressing redundant or less informative features. Finally, a pyramid-structured deep cross-attention fusion module captures latent emotional representations at multiple feature scales to improve cross-modal feature interaction. Main results. Experimental results on benchmark emotion recognition datasets demonstrated that the proposed HMCFNet achieved superior classification performance compared with several representative methods. The ablation experiments further showed the effectiveness of modality-specific feature refinement, joint cross-attention, and pyramid-structured deep fusion in improving emotion recognition performance. Significance. The proposed framework provides an effective solution for non-invasive multimodal emotion recognition by exploiting complementary information between facial expressions and PPSs. This study may facilitate the development of practical and scalable affective computing systems.

Comments (0)

No login
gif