多模态对话情感分析的目标是识别对话中每个句子的情感。现有方法的模态融合方式较简单,无法充分捕捉和利用不同模态的特性和信息。此外,这些方法更侧重于局部上下文的捕捉,特别是在处理较长对话时,往往忽略了发言者之间远距离情感信息的整合。为了解决这些问题,提出了一种基于多模态双向融合的图神经网络(Graph Neural Network Based on Multimodal Bidirectional Fusion,GMBF),该网络由多模态融合模块和远距离情感融合模块组成。多模态融合模块由三个双向融合模块组成,双向融合模块从正向和逆向两个方向融合多模态信息,通过逐步融合模态信息以确保信息的充分融合;远距离情感融合模块首先构建对话的句子信息,然后捕捉远距离发言者信息,并将其融入句子信息中,从而使模型能够更好地理解全局情感背景。实验结果表明,所提出的方法在多模态对话情感分析任务中表现优异,展现了其在多模态信息融合和全局信息提取方面的优势。
Abstract
The goal of multimodal conversational emotion recognition is to identify the emotion of each utterance in a conversation. Existing methods for modality fusion, which are relatively simple, fail to fully capture and leverage the characteristics and information of different modalities. Furthermore, these methods primarily focus on capturing local context, and often overlook the integration of long-range emotional information between speakers, especially in handling lengthy conversations. To address these issues, this paper proposes a Graph Neural Network Based on Multimodal Bidirectional Fusion (GMBF), which consists of a multimodal fusion module and a long-range emotion fusion module. The multimodal fusion module is composed of three bidirectional fusion submodules, each of which integrates multimodal information from both forward and reverse directions, ensuring comprehensive information fusion through progressive integration. The long-range emotion fusion module first constructs sentence-level information for the conversation, then captures long-range speaker information and incorporates it into the sentence representations, enabling the model to better understand the global emotional context. Experimental results demonstrate that the proposed method achieves superior performance on multimodal conversational emotion recognition tasks, demonstrating its advantages in multimodal information fusion and global information extraction.
随着社交媒体和即时通讯工具的普及,人们之间的互动变得前所未有的频繁和复杂。对话情感分析(Emotion Recognition in Conversations, ERC)作为自然语言处理(Natural Language Processing, NLP)的一个重要分支,旨在识别和理解对话中表达的情感。这一任务不仅在情感计算、心理健康评估和人机交互等领域具有广泛的应用,而且在提升用户体验和社交媒体内容分析的准确性方面发挥着至关重要的作用。
多模态信息的互补性,能够提升信息的完整性和准确性。双向长短期记忆(Bidirectional Long Short-Term Memory, BiLSTM)网络在NLP中的成功应用,证明了双向处理能够全面捕捉上下文信息。因此,本文提出了一种基于多模态双向融合的图神经网络(Graph Neural Network Based on Multimodal Bidirectional Fusion,GMBF),通过引入全局信息和建模发言者之间的互动关系改进了图神经网络,有效弥补了序列模型在全局情感捕捉上的不足[6-7]。此外,引入外部常识知识对发言者心理状态进行建模,以增强图建模中边的表达,并全面反映对话中的情感动态和参与者互动[8]。GMBF包括多模态融合模块和远距离情感融合模块,构成多模态融合模块的双向融合模块通过正向和逆向两个不同的方向融合多模态信息,确保了各个模态之间的互补信息都能被充分捕捉到,弥补单一方向融合可能遗漏的信息,提升整体信息质量;同时,远距离情感融合模块将与当前句子同一发言者的后续句子的情感信息融合到当前句子的表示中,使模型直接将远距离的相关信息聚合到当前句子中,避免了因多发言者信息融合导致的情感混淆,提升了模型的情感分析性能。本文的主要贡献如下:
GEETHAA V, MALAT, PRIYANKAD, et al. Multimodal Emotion Recognition with deep learning: Advancements, challenges, and future directions[EB/OL].[2024-05-30]. DOI: 10.1016/j.inffus.2023.102218 .
[2]
JOSHIA, BHATA, JAINA, et al. COGMEN: COntextualized GNN based Multimodal Emotion recognitioN[C]//Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Stroudsburg: Association for Computational Linguistics, 2022: 4148-4164. DOI: 10.18653/v1/2022.naacl-main.306 .
[3]
LIANZ, LIUB, TAOJ H. CTNet: Conversational transformer network for emotion recognition[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 29: 985-1000. DOI: 10.1109/TASLP.2021.3049898 .
[4]
MAH, WANGJ, LINH, et al. A multi-view network for real-time emotion recognition in conversations[EB/OL].[2022-01-25]. DOI: 10.1016/j.knosys.2021.107751 .
[5]
GHOSALD, MAJUMDERN, PORIAS, et al. DialogueGCN: A graph convolutional neural network for emotion recognition in conversation[EB/OL]. 2019: 1908.11540. DOI: 10.18653/v1/d19-1015 .
[6]
HOUG Y, SHENY L, ZHANGW Q, et al. Enhancing emotion recognition in conversation via multi-view feature alignment and memorization[C]//Findings of the Association for Computational Linguistics: EMNLP 2023. Stroudsburg: Association for Computational Linguistics, 2023: 12651-12663. DOI: 10.18653/v1/2023.findings-emnlp.842 .
[7]
WANGB, DONGG, ZHAOY, et al. Hierarchically stacked graph convolution for emotion recognition in conversation[EB/OL].[2023-03-05]. DOI: 10.1016/j.knosys.2023.110285 .
[8]
LIJ N, LINZ, FUP, et al. Past, present, and future: Conversational emotion recognition through structural modeling of psychological knowledge[C]//Findings of the Association for Computational Linguistics: EMNLP 2021. Stroudsburg: Association for Computational Linguistics, 2021: 1204-1214. DOI: 10.18653/v1/2021.findings-emnlp.104 .
HUANGL D, HUH J, CHENJ Y, et al. Aspect-based sentiment analysis method based on knowledge enhancement and multi-layer attention mechanism[J]. Journal of Wuhan University (Natural Science Edition), 2024, 70(4): 473-481. DOI: 10.14188/j.1671-8836.2023.0014(Ch ).
[11]
ZHANGX H, LIY. A cross-modality context fusion and semantic refinement network for emotion recognition in conversation[C]//Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Stroudsburg: Association for Computational Linguistics, 2023: 13099-13110. DOI: 10.18653/v1/2023.acl-long.732 .
[12]
LUN, HANZ, HANM, et al. Bi-stream graph learning based multimodal fusion for emotion recognition in conversation[EB/OL].[2024-06-30]. DOI: 10.2139/ssrn.4614720 .
[13]
ZHANGX H, CUIW G, HUB, et al. A multi-level alignment and cross-modal unified semantic graph refinement network for conversational emotion recognition[J]. IEEE Transactions on Affective Computing, 2024, 15(3): 1553-1566. DOI: 10.1109/TAFFC.2024.3354382 .
[14]
PORIAS, CAMBRIAE, HAZARIKAD, et al. Context-dependent sentiment analysis in user-generated videos[C]//Proceedings of the 55th Annual Meeting of the Association forComputational Linguistics (Volume 1: Long Papers). Stroudsburg: Association for Computational Linguistics, 2017: 873-883. DOI: 10.18653/v1/p17-1081 .
[15]
HAZARIKAD, PORIAS, MIHALCEAR, et al. ICON: Interactive conversational memory network for multimodal emotion detection[C]//Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Stroudsburg: Association for Computational Linguistics, 2018: 2594-2604. DOI: 10.18653/v1/d18-1280 .
[16]
SHENW Z, WUS Y, YANGY Y, et al. Directed acyclic graph network for conversational emotion recognition[EB/OL]. 2021: 2105.12907.DOI: 10.18653/v1/2021.acl-long.123 .
[17]
LIJ, WANGX P, LVG Q, et al. GA2MIF: Graph and attention based two-stage multi-source information fusion for conversational emotion detection[J]. IEEE Transactions on Affective Computing, 2024, 15(1): 130-143. DOI: 10.1109/TAFFC.2023.3261279 .
[18]
VASWANIA, SHAZEERN M, PARMARN, et al. Attention is all you need[EB/OL].[2017-06-12]. DOI: 10.1007/978-3-031-84300-6_13 .
[19]
BOSSELUTA, RASHKINH, SAP M, et al. COMET: Commonsense transformers for automatic knowledge graph construction[EB/OL]. 2019: 1906.05317. DOI: 10.18653/v1/p19-1470 .
[20]
SHIY S, HUANGZ J, FENGS K, et al. Masked label prediction: Unified message passing model for semi-supervised classification[EB/OL]. 2020: 2009.03509. DOI: 10.24963/ijcai.2021/214 .
[21]
CHENJ J, HOUH X, GAOJ, et al. RGCN: Recurrent Graph Convolutional Networks for Target-Dependent Sentiment Analysis[M]//Knowledge Science, Engnieering and Management. Cham: Springer International Publishing, 2019: 667-675. DOI: 10.1007/978-3-030-29551-6_59 .
[22]
JIANGJ F, WANGA, AIZAWAA. Attention-based relational graph convolutional network for target-oriented opinion words extraction[C]//Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. Stroudsburg: Association for Computational Linguistics, 2021: 1986-1977. DOI: 10.18653/v1/2021.eacl-main.170 .
[23]
BUSSOC, BULUTM, LEEC C, et al. IEMOCAP: Interactive emotional dyadic motion capture database[J]. Language Resources and Evaluation, 2008, 42(4): 335-359. DOI: 10.1007/s10579-008-9076-6 .
[24]
PORIAS, HAZARIKAD, MAJUMDERN, et al. MELD: A multimodal multi-party dataset for emotion recognition in conversations[EB/OL]. 2018: 1810.02508.DOI: 10.18653/v1/p19-1050 .
[25]
MAJUMDERN, PORIAS, HAZARIKAD, et al. DialogueRNN: An attentive RNN for emotion detection in conversations[DB/OL]. [2019-07-17]. DOI: 10.1609/aaai.v33i01.33016818 .
[26]
HUD, WEIL W, HUAIX Y. DialogueCRN: Contextual reasoning networks for emotion recognition in conversations[EB/OL]. 2021: 2106.01978. DOI: 10.18653/v1/2021.acl-long.547 .
[27]
HUD, BAOY N, WEIL W, et al. Supervised adversarial contrastive learning for emotion recognition in conversations[EB/OL]. 2023: 2306.01505. DOI: 10.18653/v1/2023.acl-long.606 .
[28]
GHOSALD, MAJUMDERN, GELBUKHA, et al. COSMIC: COmmonSense knowledge for eMotion Identification in Conversations[EB/OL]. 2020: 2010.02795. DOI: 10.18653/v1/2020.findings-emnlp.224 .
[29]
XUY, YANGM. MCM-CSD: Multi-granularity context modeling with contrastive speaker detection for emotion recognition in real-time conversation[C]//ICASSP 2024―2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). New York: IEEE Press, 2024: 11956-11960. DOI: 10.1109/ICASSP48485.2024.10446410 .
[30]
LIUY J, LIJ, WANGX P, et al. EmotionIC: Emotional inertia and contagion-driven dependency modeling for emotion recognition in conversation[J]. Science China Information Sciences, 2024, 67(8): 182103. DOI: 10.1007/s11432-023-3908-6 .
[31]
LIJ, WANGX P, LVG Q, et al. GraphCFC: A directed graph based cross-modal feature complementation approach for multimodal conversational emotion recognition[J]. IEEE Transactions on Multimedia, 2023, 26: 77-89. DOI: 10.1109/TMM.2023.3260635 .