On April 4, 2025, the latest collaborative research by the School of Psychology at Shanghai Jiao Tong University, Baidu, Midea, and East China Normal University, titled "Towards Speaker-Unknown Emotion Recognition in Conversation Via Progressive Contrastive Deep Supervision," was published online in Early Access in IEEE Transactions on Affective Computing.
Emotion recognition in conversation (ERC) has attracted increasing attention for its ability to perceive users' emotions in practical conversational applications. Most existing studies leverage speaker information based on gold-standard speaker labels to handle conversational utterances spoken alternately by different speakers. This work challenges the existing paradigm of relying on available speaker labels and considers a more realistic scenario, in which the speaker identity of each utterance is unknown during inference. The study proposes Progressive Contrastive Deep Supervision for multimodal emotion recognition in conversation (PCDS), incorporating speaker diarization and emotion recognition into a unified framework. To facilitate joint-task learning, speaker and emotion biases are progressively injected through contrastive deep supervision, with task-irrelevant contrast serving as an intermediate transition. To obtain explicit speaker dependencies, a Speaker Contrast and Clustering (SCC) module is proposed, enabling the network to partition speakers into groups even when neither speaker labels nor the number of speakers is known a priori.
Figure 1. PCDS architecture.
Research Motivation
Emotion recognition in conversation (ERC) is of substantial value in practical conversational applications, yet most existing studies rely on known speaker labels, which are often unavailable in practice. To address this issue, the study proposes a new method for emotion recognition when speaker identities are unknown. The motivation arises from practical application scenarios in which emotion must be recognized effectively without prior knowledge of speaker identity, thereby improving conversational system performance and user experience.
Research Contributions
The study proposes a Progressive Contrastive Deep Supervision (PCDS) framework that integrates speaker diarization and emotion recognition into a unified framework. By progressively applying deep supervision at different levels, PCDS can effectively model speaker and emotion representations and reconcile the underlying conflicts between them. In addition, the study introduces a Speaker Contrast and Clustering (SCC) module for multimodal speaker diarization, which clusters speakers without speaker labels and explicitly models speaker dependencies. Experimental results show that PCDS achieves state-of-the-art performance on the IEMOCAP and MELD multimodal conversation datasets.
Figure 2. Four supervision frameworks. LCE denotes cross-entropy loss and LC denotes contrastive loss. PCDS applies the corresponding task-based contrastive losses to intermediate layers.
Research Innovations
The study introduces a progressive contrastive deep supervision method that progressively injects task-specific biases and task-irrelevant contrastive losses at different levels, effectively strengthening the network's feature representations. It also designs a Speaker Contrast and Clustering (SCC) module for multimodal speaker diarization, combining audio-query fusion and cross-attention to cluster unknown speakers. This innovation addresses the challenge of unknown speakers and provides new ideas and technical approaches for multimodal emotion recognition.
Figure 3. Speaker Contrast and Clustering (SCC) module for speaker clustering, together with the speaker-aware encoder used to model speaker information.
Conclusion
In this study, a Progressive Contrastive Deep Supervision (PCDS) framework is proposed to successfully address the challenge of emotion recognition when speaker identities are unknown. By progressively applying deep supervision at different levels, PCDS effectively reconciles the underlying tension between speaker diarization and emotion recognition, while the SCC module enables clustering of unknown speakers. Experimental results show that PCDS achieves state-of-the-art performance on two multimodal conversation datasets. We hope this work can provide new ideas and methods for the future development of emotion recognition in conversation.
This study also represents the culmination of a series of related works, beginning with the LGCCT gated multimodal speech emotion recognition method (https://doi.org/10.3390/e24071010), followed by explorations of speech emotion recognition with temporal shift (https://doi.org/10.34133/icomputing.0073), and then fine-grained speech emotion recognition (https://doi.org/10.1109/ICASSP48485.2024.10446974), ultimately leading to the present study. Across this series, emotion recognition methods have been progressively refined for different contexts, laying a solid foundation for multimodal emotion recognition when speakers are unknown and enabling future behavioral experiments that further integrate psychology and AI.
Research Assistant Professor Feng Liu and Professor Aimin Zhou are co-corresponding authors. First author Siyuan Shen was an early member of Research Assistant Professor Feng Liu's affective computing team. Baidu and the School of Psychology, Shanghai Jiao Tong University, are listed as joint first affiliations. The project was supported by the Science and Technology Commission of Shanghai Municipality (Grant No. 22511105901), the National Natural Science Foundation of China (Grant No. 32471151), the National Key R&D Program of China, Key Special Project on "Proactive Health and Technological Responses to Population Aging" (Grant No. 2024YFC3606802), and the Beijing Key Laboratory of Behavior and Mental Health at Peking University.
Paper: https://ieeexplore.ieee.org/document/10949847/
Citation: S. Shen, F. Liu, H. Wang and A. Zhou, "Towards Speaker-Unknown Emotion Recognition in Conversation Via Progressive Contrastive Deep Supervision," in IEEE Transactions on Affective Computing, doi: 10.1109/TAFFC.2025.3558222.