融合数据增强与双重图对比学习的多模态讽刺检测方法

朱小栋, 刘逍遥, 江龙, 张停

小型微型计算机系统 ›› 2026, Vol. 47 ›› Issue (9) : 2182 -2191.

小型微型计算机系统 ›› 2026, Vol. 47 ›› Issue (9) : 2182 -2191. DOI: 10.20009/j.cnki.21-1106/TP.2025-0320
算法理论与人工智能

融合数据增强与双重图对比学习的多模态讽刺检测方法

    朱小栋1,2, 刘逍遥1, 江龙1, 张停1
作者信息 +

Multimodal Sarcasm Detection Method Integrating Data Augmentation and Dual Graph Contrastive Learning

    ZHU Xiaodong1,2, LIU Xiaoyao1, JIANG Long1, ZHANG Ting1
Author information +
文章历史 +

摘要

多模态讽刺检测对理解社交媒体情感倾向具有重要作用,针对该任务中跨模态交互建模不足与噪声鲁棒性弱的问题,该研究提出一种融合跨模态数据增强与双重图对比学习的模型(CM-DGCL).该方法首先利用CLIP(Contrastive Language-Image Pre-training)预训练模型提取文本与图像的深层特征,并结合OCR(Optical Character Recognition)技术识别图像文本信息;通过扰动与文本高度相关的图像块生成负样本以增强语义输入丰富性.进而设计多特征融合模块,借助双向交叉注意力机制实现文本与图像特征的细粒度交互,捕捉模态间不一致性与潜在语义关联.进一步引入双重图对比学习策略,以原始图文对为正样本、扰动图文对为负样本,通过对比学习损失约束语义一致性并动态优化相似度分布,从而提升模型对噪声的鲁棒性及对细微讽刺线索的感知能力.在两个公开的多模态讽刺数据集上的实验表明,CM-DGCL相较当前最优模型,准确率分别提升了2.27%和3.25%,F1值分别提升了2.56%和3.68%,显著优于现有方法.

Abstract

Multimodal sarcasm detection is essential for understanding affective tendencies in social media.To address the issues of insufficient cross-modal interaction modeling and weak noise robustness in multimodal sarcasm detection tasks,this study proposed a model called CM-DGCL (Cross-modal Data augmentation and Dual-graph contrastive Learning).The method first utilized the CLIP (Contrastive Language-Image Pre-training) pre-trained model to extract deep features from text and images,and combined OCR (Optical Character Recognition) technology to recognize textual information within images.Negative samples were generated by perturbing image patches highly correlated with the text to enhance semantic input richness.Subsequently,the method designed a multi-feature fusion module,leveraging a bidirectional cross-attention mechanism to achieve fine-grained interaction between text and image features,capturing inter-modal inconsistencies and potential semantic relationships.Further,the method introduced a dual-graph contrastive learning strategy.This strategy treated the original image-text pair as a positive sample and the perturbed pair as a negative sample.The contrastive learning loss constrained semantic consistency and dynamically optimized the similarity distribution,thereby enhancing the model′s robustness to noise and its ability to perceive subtle sarcastic cues.Experimental results on two public multimodal sarcasm detection datasets demonstrate that CM-DGCL surpasses the previous state-of-the-art models,achieving improvements in accuracy by 2.27% and 3.25%,and in F1-score by 2.56% and 3.68%,respectively,significantly outperforming existing approaches.

关键词

多模态讽刺检测 / 数据增强 / 对比学习 / 交叉注意力 / 预训练模型

Key words

multimodal sarcasm detection / data augmentation / contrastive learning / cross-attention / pre-trained model

引用本文

引用格式 ▾
朱小栋, 刘逍遥, 江龙, 张停. 融合数据增强与双重图对比学习的多模态讽刺检测方法[J]. 小型微型计算机系统, 2026, 47(9): 2182-2191 DOI:10.20009/j.cnki.21-1106/TP.2025-0320

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1] Zhang D,Li S,Zhu Q,et al.Effective sentiment-relevant word selection for multi-modal sentiment analysis in spoken language[C]//Proceedings of the 27th ACM International Conference on Multimedia,2019:148-156.
[2] Zhang D,Wei S,Li S,et al.Multi-modal graph fusion for named entity recognition with targeted visual guidance[C]//Proceedings of the AAAI Conference on Artificial Intelligence,2021:14347-14355.
[3] Joshi A,Sharma V,Bhattacharyya P.Harnessing context incongruity for sarcasm detection[C]//Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2:Short Papers),2015:757-762.
[4] Zhang M,Zhang Y,Fu G.Tweet sarcasm detection using deep neural network[C]//26th International Conference on Computational Linguistics:Technical Papers,2016:2449-2460.
[5] HAN H,ZHAO Q T,SUN T Y,et al.Contextual sarcasm detection model for social media comments[J].Computer Engineering,2021,47(1):66-71.
[6] Tay Y,Tuan L A,Hui S C,et al.reasoning with sarcasm by reading in-between[C]//56th Annual Meeting of the Association for Computational Linguistics (Volume 1:Long Papers),2018:1010-1020.
[7] Schifanella R,De Juan P,Tetreault J,et al.Detecting sarcasm in multimodal social platforms[C]//24th ACM International Conference on Multimedia,2016:1136-1145.
[8] Cai Y,Cai H,Wan X.Multi-modal sarcasm detection in twitter with hierarchical fusion model[C]//Proceedings of the 57th Annual Meeting of the Association For Computational Linguistics,2019:2506-2515.
[9] WU Y B,ZENG W S,GAO H,et al.Study on multimodal sarcasm explanation based on dual-stream residual fusion[J].Journal of Chinese Computer Systems,2024,45(11):2628-2635.
[10] Pan H,Lin Z,Fu P,et al.Modeling intra and inter-modality incongruity for multi-modal sarcasm detection[C]//Findings of the Association for Computational Linguistics,2020:1383-1392.
[11] Liu H,Wang W,Li H.Towards multi-modal sarcasm detection via hierarchical congruity modeling with knowledge enhancement[C]//Conference on Empirical Methods in Natural Language Processing,2022:4995-5006.
[12] Qiao Y,Jing L,Song X,et al.Mutual-enhanced incongruity learning network for multi-modal sarcasm detection[C]//AAAI Conference on Artificial Intelligence,2023:9507-9515.
[13] Liang B,Lou C,Li X,et al.Multi-modal sarcasm detection via cross-modal graph convolutional network[C]//60th Annual Meeting of the Association for Computational Linguistics (Volume 1:Long Papers),2022:1767-1777.
[14] YU B G,JI X H.Detecting multimodal sarcasm based on ADGCN-MFM[J].Data Analysis and Knowledge Discovery,2023,7(10):85-94.
[15] Wei Y,Yuan S,Zhou H,et al.G2SAM:graph-based global semantic awareness method for multimodal sarcasm detection[C]//AAAI Conference on Artificial Intelligence,2024:9151-9159.
[16] Gao L,Sheng N,Liu Y,et al.TCIP:network with topology capture and incongruity perception for sarcasm detection[J].Information Fusion,2025,117:102918,doi:10.1016/j.inffus.2024.102918.
[17] Tian Y,Xu N,Zhang R,et al.Dynamic routing transformer network for multimodal sarcasm detection[C]//61st Annual Meeting of the Association for Computational Linguistics (Volume 1:Long Papers),2023:2468-2480.
[18] Xi Z,Yu B,Wang H.Multimodal sarcasm detection based on sentiment-clue inconsistency global detection fusion network[J].Expert Systems with Applications,2025,275:127020,doi:10.1016/j.eswa.2025.127020.
[19] Lu Q,Long Y,Sun X,et al.Fact-sentiment incongruity combination network for multimodal sarcasm detection[J].Information Fusion,2024,104:102203,doi:10.1016/j.inffus.2023.102203.
[20] LIN J X,ZHU X D.CMHICL:multimodal sarcasm detection based on cross-modal hierarchical interaction network and contrastive learning[J].Application Research of Computers,2024,41(9):2620-2627.
[21] Mao R,Lin C,Guerin F.Word embedding and WordNet based metaphor identification and interpretation[C]//Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics,2018:1222-1231.
[22] Zhang K,Li Y,Wang J,et al.Real-time video emotion recognition based on reinforcement learning and domain knowledge[J].IEEE Transactions on Circuits and Systems for Video Technology,2021,32(3):1034-1047.
[23] Ge M,Mao R,Cambria E.Explainable metaphor identification inspired by conceptual metaphor theory[C]//AAAI Conference on Artificial Intelligence,2022:10681-10689.
[24] Li W,Zhu L,Mao R,et al.SKIER:a symbolic knowledge integrated model for conversational emotion recognition[C]//AAAI Conference on Artificial Intelligence,2023:13121-13129.
[25] Tang B,Lin B,Yan H,et al.Leveraging generative large language models with visual instruction and demonstration retrieval for multimodal sarcasm detection[C]//Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies (Volume 1:Long Papers),2024:1732-1742.
[26] Wei Y,Duan M,Zhou H,et al.Towards multimodal sarcasm detection via label-aware graph contrastive learning with back-translation augmentation[J].Knowledge-Based Systems,2024,300:112109,doi:10.1016/j.knosys.2024.112109.
[27] LUO W P,HUANG D G.Enhanced cross-modal image-text retrieval with large models[J].Journal of Chinese Computer Systems,2025,46(7):1544-1553.
[28] Yue T,Mao R,Wang H,et al.KnowleNet:knowledge fusion network for multimodal sarcasm detection[J].Information Fusion,2023,100:101921,doi:10.1016/j.inffus.2023.101921.
[29] Niu Z,Xie Z,Xu T,et al.Knowledge-enhanced multi-perspective incongruity perception network for multimodal sarcasm detection[C]//IEEE International Conference on Multimedia and Expo,2024:1-6.
[30] Zhou L,Palangi H,Zhang L,et al.Unified vision-language pre-training for image captioning and vqa[C]//Proceedings of the AAAI Conference on Artificial Intelligence,2020:13041-13049.
[31] CHEN F L,ZHANG D Z,HAN M L,et al.Vlp:a survey on vision-language pre-training[J].Machine Intelligence Research,2023,20(1):38-56.
[32] Li L H,Yatskar M,Yin D,et al.Visualbert:a simple and performant baseline for vision and language[J].arXiv Preprint,2019,arXiv:1908.03557.
[33] Radford A,Kim J W,Hallacy C,et al.Learning transferable visual models from natural language supervision[C]//International Conference on Machine Learning,2021:8748-8763.
[34] Qin L,Huang S,Chen Q,et al.MMSD2.0:towards a reliable multi-modal sarcasm detection system[C]//Findings of the Association for Computational Linguistics,2023:10834-10845.
[35] Li J,Li D,Savarese S,et al.Blip-2:bootstrapping language-image pre-training with frozen image encoders and large language models[C]//International Conference on Machine Learning,2023:19730-19742.
[36] CHEN Y Y,MA J.Detecting multimodal sarcasm based on SC-attention mechanism[J].Data Analysis and Knowledge Discovery,2022,6(9):40-51.
[37] Liu Y,Ott M,Goyal N,et al.Roberta:a robustly optimized bert pretraining approach[J].arXiv:1907.11692,2019.
[38] Cubuk E D,Zoph B,Shlens J,et al.Randaugment:practical automated data augmentation with a reduced search space[C]//IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops,2020:702-703.
[39] Chen Y,Shi S,Huang H.Which is more faithful,seeing or saying? Multimodal sarcasm detection exploiting contrasting sentiment knowledge[J].CAAI Transactions on Intelligence Technology,2025,10(2):375-386.
[40] Liang B,Luo W,Li X,et al.Enhancing aspect-based sentiment analysis with supervised contrastive learning[C]//30th ACM International Conference on Information & Knowledge Management,2021:3242-3247.
[41] YAN Y,YANG D,YIN D C.Contrastive-based prompt-tuning sentiment analysis method incorporating large language model knowledge[J].Journal of Intelligence,2023,42(11):126-134.
[42] CHEN N,LIU F,DONG C W,et al.Few-shot image classification based on local contrastive learning and novel class feature generation[J].Pattern Recognition and Artificial Intelligence,2024,37(10):936-946.
[43] QIAN Z S,HUANG H,WAN Z L.The multi-behavior graph contrastive learning recommendation method with self-attention mechanism[J].Acta Electronica Sinica,2024,52(11):3684-3698.
[44] Yang Z,Lin C,Qin Y,et al.JGC-IAGCL:fusing joint graph convolution and intent-aware graph contrastive learning for explainable recommendation[J].Information Fusion,2025,123:103258,doi:10.1016/j.inffus.2025.103258.
[45] Liang B,Gui L,He Y,et al.Fusion and discrimination:a multimodal graph contrastive learning framework for multimodal sarcasm detection[J].IEEE Transactions on Affective Computing,2024,15(4):1874-1888.
[46] Wu Q,Fang W,Zhong W,et al.Dual-level adaptive incongruity-enhanced model for multimodal sarcasm detection[J].Neurocomputing,2025,612:128689,doi:10.1016/j.neucom.2024.128689.
[47] Qiao Y,Jing L,Song X,et al.Mutual-enhanced incongruity learning network for multi-modal sarcasm detection[C]//AAAI Conference on Artificial Intelligence,2023:9507-9515.
[48] Long X,Gan C,De Melo G,et al.Multimodal keyless attention fusion for video classification[C]//32nd AAAI Conference on Artificial Intelligence,2018:7202-7209.
[49] Kim Y.Convolutional neural networks for sentence classification[J].arXiv:1408.5882,2014.
[50] Graves A,Schmidhuber J.Framewise phoneme classification with bidirectional LSTM and other neural network architectures[J].Neural Networks,2005,18(5-6):602-610.
[51] Xiong T,Zhang P,Zhu H,et al.Sarcasm detection with self-matching networks and low-rank bilinear pooling[C]//The World Wide Web Conference,2019:2115-2124.
[52] He K,Zhang X,Ren S,et al.Deep residual learning for image recognition[C]//IEEE Conference on Computer Vision and Pattern Recognition,2016:770-778.
[53] Dosovitskiy A,Beyer L,Kolesnikov A,et al.An image is worth 16×16 words:transformers for image recognition at scale[J].arXiv:2010.11929,2020.
[54] Wen C,Jia G,Yang J.Dip:dual incongruity perceiving network for sarcasm detection[C]//IEEE/CVF Conference on Computer Vision and Pattern Recognition,2023:2540-2550.
[55] Chen Z,Lin H,Luo Z,et al.CofiPara:a coarse-to-fine paradigm for multimodal sarcasm target identification with large multimodal models[C]//62nd Annual Meeting of the Association for Computational Linguistics (Volume 1:Long Papers),2024:9663-9687.
[56] Zhu Z,Zhuang X,Zhang Y,et al.TFCD:towards multi-modal sarcasm detection via training-free counterfactual debiasing[C]//33rd International Joint Conference on Artificial Intelligence,2024:6687-6695.
附中文参考文献:
[5] 韩 虎,赵启涛,孙天岳,等.面向社交媒体评论的上下文语境讽刺检测模型[J].计算机工程,2021,47(1):66-71.
[9] 吴运兵,曾炜森,高 航,等.基于双流残差融合的多模态讽刺解释研究[J].小型微型计算机系统,2024,45(11):2628-2635.
[14] 余本功,季晓晗.基于ADGCN-MFM的多模态讽刺检测研究[J].数据分析与知识发现,2023,7(10):85-94.
[20] 林洁霞,朱小栋.CMHICL:基于跨模态分层交互网络和对比学习的多模态讽刺检测[J].计算机应用研究,2024,41(9):2620-2627.
[27] 罗文培,黄德根.大模型增强的跨模态图文检索方法[J].小型微型计算机系统,2025,46(7):1544-1553.
[36] 陈圆圆,马 静.基于SC-Attention机制的多模态讽刺检测研究[J].数据分析与知识发现,2022,6(9):40-51.
[41] 严 豫,杨 笛,尹德春.融合大语言模型知识的对比提示情感分析方法[J].情报杂志,2023,42(11):126-134.
[42] 陈 宁,刘 凡,董晨炜,等.基于局部对比学习与新类特征生成的小样本图像分类[J].模式识别与人工智能,2024,37(10):936-946.
[43] 钱忠胜,黄 恒,万子珑.融合自注意力机制的多行为图对比学习推荐方法[J].电子学报,2024,52(11):3684-3698.

基金资助

国家自然科学基金面上项目(71871144)资助.

AI Summary AI Mindmap

0

访问

0

被引

详细

导航
相关文章

AI思维导图

/