跨模态自适应融合网络用于多模态关系抽取

侯长波 ,  王睿奇 ,  韩吉南 ,  汤曜辉

电子科技大学学报 ›› 2026, Vol. 55 ›› Issue (4) : 590 -599.

PDF (931KB)
电子科技大学学报 ›› 2026, Vol. 55 ›› Issue (4) : 590 -599. DOI: 10.12178/1001-0548.2025165
计算机工程与应用

跨模态自适应融合网络用于多模态关系抽取

作者信息 +

Feature-adaptive fusion network for multimodal relation extraction

Author information +
文章历史 +
PDF (953K)

摘要

针对现有多模态关系抽取(MRE)模型难以有效解决多模态特征融合不充分和实际应用下多模态数据不相关的问题,提出一种跨模态自适应融合网络(CAFNeT),旨在提升多模态关系抽取的性能。首先基于扩散生成模型生成新的图像数据,并与原始文本数据组成新的图文对作为网络的辅助信息;其次在多层次跨模态交互模块中设计了一种自适应多头交叉注意力机制,分别从句子层面和词汇层面实现粗细粒度的跨模态特征交互,深入地捕捉模态间的语义依赖关系;最后,在多模态自适应融合模块中设计了一种基于辅助信息的图文不相关修正机制,通过自适应调节机制动态调整原始图文对信息与辅助信息在最终表示中的贡献比例,有效抑制图文不相关数据信息的影响,同时避免关键信息丢失。在公开多模态关系抽取数据集上的实验结果表明,CAFNeT在精确率、召回率和F1值上分别达到91.30%、90.16%与90.72%,相较于基准模型均有一定提升,进一步的消融研究证实了所提方法的有效性。

Abstract

To address the challenges of insufficient multimodal feature fusion and irrelevant multimodal data in practical applications faced by existing multimodal relation extraction (MRE) models, this paper proposes a cross-modal adaptive fusion network (CAFNeT), aiming to enhance MRE performance. Firstly, novel image data is generated using diffusion generative models and combined with the original text data to form new image-text pairs serving as auxiliary information for the network. Secondly, a multi-level cross-modal interaction module incorporating an adaptive multi-head cross-attention mechanism is designed. This mechanism facilitates coarse- and fine-grained cross-modal feature interactions at both the sentence and word levels, enabling the deep capture of semantic dependencies between modalities. Finally, a multimodal adaptive fusion module is introduced, featuring an image-text irrelevance correction mechanism based on auxiliary information. This mechanism dynamically adjusts the contribution proportions of the original image-text pair information and the auxiliary information in the final representation through an adaptive regulation mechanism. This effectively suppresses the influence of irrelevant image-text data while avoiding the loss of critical information. Experimental results on public MRE datasets demonstrate that CAFNeT achieves precision, recall, and F1-scores of 91.30%, 90.16%, and 90.72%, respectively, showing improvements over baseline models. Further ablation studies confirm the effectiveness of the proposed method.

关键词

关系抽取 / 多模态信息 / 交叉注意力机制 / 跨模态融合 / 自适应权重 / 扩散模型

Key words

relation extraction / multimodal information / cross-attention mechanism / cross-modal fusion / adaptive weight / diffusion model

引用本文

引用格式 ▾
侯长波,王睿奇,韩吉南,汤曜辉. 跨模态自适应融合网络用于多模态关系抽取[J]. 电子科技大学学报, 2026, 55(4): 590-599 DOI:10.12178/1001-0548.2025165

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

Zheng Changmeng, Feng Junhao, Fu Ze, et al. Multimodal relation extraction with efficient graph alignment[C]//Proceedings of the 29th ACM International Conference on Multimedia. New York: ACM, 2021: 5298-5306.

[2]

Chen Xiang, Zhang Ningyu, Li Lei, et al. Good visual guidance make a better extractor: Hierarchical visual prefix for multimodal entity and relation extraction[C]//Proceedings of the Association for Computational Linguistics: NAACL 2022. Seattle: Association for Computational Linguistics, 2022: 1607-1618.

[3]

Xu Bo, Huang Shizhou, Du Ming, et al. A unified visual prompt tuning framework with mixture-of-experts for multimodal information extraction[C]//Proceedings of the 28th International Conference on Database Systems for Advanced Applications. Tianjin: Springer, 2023: 544-554.

[4]

Wang Min, Chen Hongbin, Shen Dingcai, et al. RSRNeT: A novel multi-modal network framework for named entity recognition and relation extraction[J].PeerJ Computer Science, 2024, 10: e1856.

[5]

Zheng Changmeng, Wu Zhiwei, Feng Junhao, et al. MNRE: A challenge multimodal dataset for neural relation extraction with visual evidence in social media posts[C]//Proceedings of the IEEE International Conference on Multimedia and Expo. Shenzhen: IEEE,2021: 1-6.

[6]

李冬梅, 张扬, 李东远, . 实体关系抽取方法研究综述[J].计算机研究与发展, 2020, 57(7): 1424-1448.

[7]

Li Dongmei, Zhang Yang, Li Dongyuan, et al. Review of entity relation extraction methods[J].Journal of Computer Research and Development, 2020, 57(7): 1424-1448. (in Chinese)

[8]

曾泽凡, 胡星辰, 成清, . 基于预训练语言模型的知识图谱研究综述[J].计算机科学, 2025, 52(1): 1-33.

[9]

Zeng Zefan, Hu Xingchen, Cheng Qing, et al. Survey of research on knowledge graph based on pre-trained language models[J].Computer Science, 2025, 52(1): 1-33. (in Chinese)

[10]

Humphreys K, Gaizauskas R, Azzam S, et al. University of Sheffield: Description of the LaSIE-II system as used for MUC-7[C]//Proceedings of the 7th Message Understanding Conference (MUC-7). Fairfax, Virginia: Association for Computational Linguistics, 1998: 84-89.

[11]

Neelakantan A, Collins M. Learning dictionaries for named entity recognition using minimal supervision[C]//Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics. Gothenburg: Association for Computational Linguistics, 2014: 452-461.

[12]

Sun Xia, Dong Lehong. Feature-based approach to Chinese term relation extraction[C]//Proceedings of the International Conference on Signal Processing Systems. Singapore: IEEE, 2009: 410-414.

[13]

郭喜跃, 何婷婷, 胡小华, . 基于句法语义特征的中文实体关系抽取[J].中文信息学报, 2014, 28(6): 183-189.

[14]

Guo Xiyue, He Tingting, Hu Xiaohua, et al. Chinese named entity relation extraction based on syntactic and semantic features[J].Journal of Chinese Information Processing, 2014, 28(6): 183-189. (in Chinese)

[15]

甘丽新, 万常选, 刘德喜, . 基于句法语义特征的中文实体关系抽取[J].计算机研究与发展, 2016, 53(2): 284-302.

[16]

Gan Lixin, Wan Changxuan, Liu Dexi, et al. Chinese named entity relation extraction based on syntactic and semantic features[J].Journal of Computer Research and Development, 2016, 53(2): 284-302. (in Chinese)

[17]

Nguyen T H, Grishman R. Relation extraction: Perspective from convolutional neural networks[C]//Proceedings of the 1st Workshop on Vector Space Modeling for Natural Language Processing. Denver: Association for Computational Linguistics, 2015: 39-48.

[18]

Zeng Daojian, Liu Kang, Lai Siwei, et al. Relation classification via convolutional deep neural network[C]//Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers. Dublin: Association for Computational Linguistics, 2014: 2335-2344.

[19]

Soares L B, Fitzgerald N, Ling J, et al. Matching the blanks: Distributional similarity for relation learning[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence: Association for Computational Linguistics, 2019: 2895-2905.

[20]

王永胜, 李培峰, 王中卿, . 多模态信息抽取研究综述[J].软件学报, 2025, 36(4): 1665-1691.

[21]

Wang Yongsheng, Li Peifeng, Wang Zhongqing, et al. Survey on multimodal information extraction research[J].Journal of Software, 2025, 36(4): 1665-1691. (in Chinese)

[22]

Wang Xinyu, Cai Jiong, Jiang Yong, et al. Named entity and relation extraction with multi-modal retrieval[C]//Proceedings of the Association for Computational Linguistics: EMNLP 2022. Abu Dhabi: Association for Computational Linguistics, 2022: 5925-5936.

[23]

Hu Xuming, Chen Junzhe, Liu Aiwei, et al. Prompt me up: Unleashing the power of alignments for multimodal entity and relation extraction[C]//Proceedings of the 31st ACM International Conference on Multimedia. Ottawa: ACM, 2023: 5185-5194.

[24]

Cui Shiyao, Cao Jiangxia, Cong Xin, et al. Enhancing multimodal entity and relation extraction with variational information bottleneck[J].IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024, 32: 1274-1285.

[25]

Wu Shengqiong, Fei Hao, Cao Yixin, et al. Information screening whilst exploiting! Multimodal relation extraction with feature denoising and multimodal topic modeling[C]//Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Toronto: Association for Computational Linguistics, 2023: 14734-14751.

[26]

Gong Yunchao, Lv Xueqiang, Yuan Zhu, et al. CE-DCVSI: Multimodal relational extraction based on collaborative enhancement of dual-channel visual semantic information[J].Expert Systems with Applications, 2025, 262: 125608.

[27]

He Xinyu, Li Shixin, Zhang Yuning, et al. The more quality information the better: Hierarchical generation of multi-evidence alignment and fusion model for multimodal entity and relation extraction[J].Information Processing & Management, 2025, 62(1): 103875.

[28]

Zheng Changmeng, Feng Junhao, Cai Yi, et al. Rethinking multimodal entity and relation extraction from a translation point of view[C]//Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. Toronto: Association for Computational Linguistics, 2023: 6810-6824.

[29]

Schuhmann C, Vencu R, Beaumont R, et al. LAION-400M: Open dataset of CLIP-filtered 400 million image-text pairs[C]//Proceedings of the 35th International Conference on Neural Information Processing Systems. Sydney: Curran Associates Inc., 2021.

[30]

Rombach R, Blattmann A, Lorenz D, et al. High-resolution image synthesis with latent diffusion models[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE,2022: 10674-10685.

[31]

Zhang Dong, Wei Suzhong, Li Shoushan, et al. Multi-modal graph fusion for named entity recognition with targeted visual guidance[C]//Proceedings of the 35th AAAI Conference on Artificial Intelligence. [S.l.]: AAAI Press,2021: 14347-14355.

[32]

Yang Zhengyuan, Gong Boqing, Wang Liwei, et al. A fast and accurate one-stage approach to visual grounding[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Seoul: IEEE, 2019: 4682-4692.

[33]

Devlin J, Chang Mingwei, Lee K, et al. BERT: Pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Minneapolis: Association for Computational Linguistics, 2019: 4171-4186.

[34]

He Kaiming, Zhang Xiangyu, Ren Shaoqing, et al. Deep residual learning for image recognition[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016: 770-778.

[35]

Lu Di, Neves L, Carvalho V, et al. Visual attention model for name tagging in multimodal social media[C]//Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics. Melbourne: Association for Computational Linguistics, 2018: 1990-1999.

[36]

Zhang Qi, Fu Jinlan, Liu Xiaoyu, et al. Adaptive co-attention network for named entity recognition in tweets[C]//Proceedings of the 32nd AAAI Conference on Artificial Intelligence. New Orleans: AAAI, 2018: 5674-5681.

[37]

Zeng Daojian, Liu Kang, Chen Yubo, et al. Distant supervision for relation extraction via piecewise convolutional neural networks[C]//Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Lisbon: Association for Computational Linguistics, 2015: 1753-1762.

[38]

Zhong Zexuan, Chen Danqi. A frustratingly easy approach for entity and relation extraction[C]//Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT 2021). [S.l.]: Association for Computational Linguistics,2021: 50-61.

[39]

Chen Xiang, Zhang Ningyu, Li Lei, et al. Hybrid transformer with multi-level fusion for multimodal knowledge graph completion[C]//Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. Madrid: ACM, 2022: 904-915.

[40]

Dai Yinglong, Gao Feng, Zeng Daojian. An alignment and matching network with hierarchical visual features for multimodal named entity and relation extraction[C]//Proceedings of the 30th International Conference on Neural Information Processing. Changsha: Springer, 2023: 298-310.

[41]

Ding Ming, Yang Zhuoyi, Hong Wenyi, et al. CogView: Mastering text-to-image generation via transformers[C]//Proceedings of the 35th International Conference on Neural Information Processing Systems. Red Hook: Curran Associates Inc., 2021: 19822-19835.

[42]

Zhou Yufan, Zhang Ruiyi, Chen Changyou, et al. Towards language-free training for text-to-image generation[C]//Proceedings of 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans: IEEE,2022: 17886-17896.

[43]

Heusel M, Ramsauer H, Unterthiner T, et al. GANs trained by a two time-scale update rule converge to a local Nash equilibrium[C]//Proceedings of the 31st International Conference on Neural Information Processing Systems. Long Beach: Curran Associates Inc., 2017: 6629-6640.

[44]

Salimans T, Goodfellow I, Zaremba W, et al. Improved techniques for training GANs[C]//Proceedings of the 30th International Conference on Neural Information Processing Systems. Barcelona: Curran Associates Inc., 2016: 2234-2242.

[45]

Szegedy C, Vanhoucke V, Ioffe S, et al. Rethinking the inception architecture for computer vision[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016: 2818-2826.

基金资助

中央高校基本科研业务费专项资金(3072025ZN0801)

AI Summary AI Mindmap
PDF (931KB)

57

访问

0

被引

详细

导航
相关文章

AI思维导图

/