基于多源知识跨模态对齐与融合的医学报告生成方法

张晓丹 ,  王博岳 ,  王锡文 ,  孙善斌 ,  李晓理 ,  刘兆会

北京工业大学学报 ›› 2026, Vol. 52 ›› Issue (8) : 858 -869.

PDF (3294KB)
北京工业大学学报 ›› 2026, Vol. 52 ›› Issue (8) : 858 -869. DOI: 10.11936/bjutxb2024090012
研究论文

基于多源知识跨模态对齐与融合的医学报告生成方法

作者信息 +

Medical Report Generation Method Based on Multi-source Knowledge Cross-modal Alignment and Fusion

Author information +
文章历史 +
PDF (3372K)

摘要

针对医学影像和报告在数据内容和数据分布2个层面存在的数据偏差问题,提出一种基于多源知识跨模态对齐与融合的医学报告生成方法(medical report generation method based on multi-source knowledge cross-modal alignment and fusion, MRGM-MKCAF)。首先,将放射学相关的医学概念作为先验知识,并将案例影像和报告作为经验知识,构建多模态知识库;然后,提出一种基于多头注意力机制的跨模态对齐方法,利用多个注意力头分别关注知识库中多种模态特征的子空间,并进行跨模态细粒度特征对齐,获得更加全面准确的编码特征;最后,提出一种基于门控机制的解码模块,对多源编码特征进行动态的特征选取,去除冗余和无关信息,进而用于报告生成。在2组数据集上的实验结果表明,提出的方法显著提高了生成报告的准确性、流畅性和完整性。

Abstract

To address data bias in medical imaging and reporting at both the data content and distribution, a medical report generation method based on multi-source knowledge cross-modal alignment and fusion (MRGM-MKCAF) is proposed. First, radiology-related medical concepts were used as prior knowledge, and case images and reports were used as empirical knowledge to construct a multi-modal knowledge base. Subsequently, a cross-modal alignment method employing multi-head attention mechanism was proposed, which used multiple attention heads to focus on the subspaces of various modal features in the knowledge base and performed fine-grained cross-modal feature alignment to obtain more comprehensive and accurate encoded features. Finally, a decoding module based on gating mechanism was proposed to dynamically select features from multi-source encoded features, removing redundancy and irrelevant information, which was then used for report generation. Experimental results on two datasets show that the proposed method significantly enhances the accuracy, fluency, and completeness of generated reports.

关键词

医学报告生成 / 多源知识 / 跨模态对齐 / 跨模态融合 / 注意力机制 / 医学影像

Key words

medical report generation / multi-source knowledge / cross-modal alignment / cross-modal fusion / attention mechanism / medical imaging

引用本文

引用格式 ▾
张晓丹,王博岳,王锡文,孙善斌,李晓理,刘兆会. 基于多源知识跨模态对齐与融合的医学报告生成方法[J]. 北京工业大学学报, 2026, 52(8): 858-869 DOI:10.11936/bjutxb2024090012

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

Johnson A E W, Pollard T J, Greenbaum N R, et al. MIMIC—CXR—JPG, a large publicly available database of labeled chest radiographs [PP/OL]. V5. arXiv (2019—11—14) [2024—03—20]. https://doi.org/10.48550/arXiv.1901.07042.

[2]

Demner—Fushman D, Kohli M D, Rosenman M B, et al. Preparing a collection of radiology examinations for distribution and retrieval [J]. Journal of the American Medical Informatics Association, 2016, 23(2): 304-310.

[3]

Yan A, He Z X, Lu X, et al. Weakly supervised contrastive learning for chest X—ray report generation[C]// Findings of the Association for Computational Linguistics: EMNLP 2021. Stroudsburg, PA: ACL, 2021: 4009-4015.

[4]

Jiang S Q, Zhu Y H, Liu C L, et al. Dataset bias in few—shot image recognition [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(1): 229-246.

[5]

Tsotsos J K, Culhane S M, Yin K W W, et al. Modeling visual attention via selective tuning [J]. Artificial Intelligence, 1995, 78(1/2): 507-545.

[6]

Ji J Z, Xu C, Zhang X D, et al. Spatio—temporal memory attention for image captioning [J]. IEEE Transactions on Image Processing, 2020, 29: 7615-7628.

[7]

Zabrodsky H, Peleg S . Attentive transmission [J]. Journal of Visual Communication and Image Representation, 1990, 1(2): 189-198.

[8]

刘茂福, 施琦, 聂礼强 . 基于视觉关联与上下文双注意力的图像描述生成方法 [J]. 软件学报, 2022, 33(9): 3210-3222.

[9]

Liu M F, Shi Q, Nie L Q . Image captioning based on visual relevance and context dual attention [J]. Journal of Software, 2022, 33(9): 3210-3222. (in Chinese)

[10]

李博涵, 向宇轩, 封顶, . 融合知识感知与双重注意力的短文本分类模型 [J]. 软件学报, 2022, 33(10): 3565-3581.

[11]

Li B H, Xiang Y X, Feng D, et al. Short text classification model combining knowledge aware and dual attention [J]. Journal of Software, 2022, 33(10): 3565-3581. (in Chinese)

[12]

Jing B Y, Xie P T, Xing E . On the automatic generation of medical imaging reports [PP/OL]. V3. arXiv (2018—07—20) [2024—03—20]. https://doi.org/10.48550/arXiv.1711.08195.

[13]

Krause J, Johnson J, Krishna R, et al. A hierarchical approach for generating descriptive image paragraphs [C]// 2017 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2017: 3337-3345.

[14]

Chen Z H, Shen Y L, Song Y, et al. Cross—modal memory networks for radiology report generation [PP/OL]. V1. arXiv (2022—04—28)[2024—03—20]. https://doi.org/10.48550/arXiv.2204.13258.

[15]

Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need [C]// Proceedings of the 31st International Conference on Neural Information Processing Systems. Red Hook, NY: Curran Associates Inc., 2017: 6000-6010.

[16]

Li X Y, Wang Z H, Yang J H, et al. KERM: knowledge enhanced reasoning for vision—and—language navigation [C]// 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2023: 2583-2592.

[17]

Nooralahzadeh F, Gonzalez N P, Frauenfelder T, et al. Progressive transformer—based generation of radiology reports [C]// Findings of the Association for Computational Linguistics: EMNLP 2021. Stroudsburg, PA: ACL, 2021: 2824-2832.

[18]

Song X, Zhang X D, Ji J Z, et al. Cross—modal contrastive attention model for medical report generation [C]// Proceedings of the 29th International Conference on Computational Linguistics. [S.l.]: ICCI, 2022: 2388-2397.

[19]

Zhang Y X, Wang X S, Xu Z Y, et al. When radiology report generation meets knowledge graph[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2020, 34(7): 12910-12917.

[20]

Liu F L, Wu X, Ge S, et al. Exploring and distilling posterior and prior knowledge for radiology report generation [C]// 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2021: 13748-13757.

[21]

Yang S X, Wu X, Ge S, et al. Knowledge matters: chest radiology report generation with general and specific knowledge [J]. Medical Image Analysis, 2022, 80: 102510.

[22]

Zhang J S, Shen X X, Wan S H, et al. A novel deep learning model for medical report generation by inter—intra information calibration [J]. IEEE Journal of Biomedical and Health Informatics, 2023, 27(10): 5110-5121.

[23]

韩琪, 张淑军, 谭立玮, . 变尺度特征融合与交叉训练的医学报告生成方法 [J]. 计算机辅助设计与图形学学报, 2024, 36(5): 795-804.

[24]

Han Q, Zhang S J, Tan L W, et al. Medical report generation method based on multi—scale feature fusion and cross—training [J]. Journal of Computer—Aided Design & Computer Graphics, 2024, 36(5): 795-804. (in Chinese)

[25]

谭立玮, 张淑军, 韩琪, . 面向医学影像报告生成的门归一化编解码网络 [J]. 智能系统学报, 2024, 19(2): 411-419.

[26]

Tan L W, Zhang S J, Han Q, et al. Gate normalized encoder—decoder network for medical image report generation [J]. CAAI Transactions on Intelligent Systems, 2024, 19(2): 411-419. (in Chinese)

[27]

Deng J, Dong W, Socher R, et al. ImageNet: a large—scale hierarchical image database [C]// 2009 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2009: 248-255.

[28]

Irvin J, Rajpurkar P, Ko M, et al. CheXpert: a large chest radiograph dataset with uncertainty labels and expert comparison [J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2019, 33(1): 590-597.

[29]

Huang G, Liu Z, Van Der Maaten L, et al. Densely connected convolutional networks [C]// 2017 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2017: 2261-2269.

[30]

Kaur N, Mittal A, Singh G . Methods for automatic generation of radiological reports of chest radiographs: a comprehensive survey [J]. Multimedia Tools and Applications, 2022, 81(10): 13409-13439.

[31]

Papineni K, Roukos S, Ward T, et al. BLEU: a method for automatic evaluation of machine translation [C]// Proceedings of the 40th Annual Meeting on Association for Computational Linguistics. Stroudsburg, PA: ACL, 2002: 311-318.

[32]

Banerjee S, Lavie A . METEOR: an automatic metric for MT evaluation with improved correlation with human judgments[C]// Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization. Stroudsburg, PA: ACL, 2005: 65-72.

[33]

Lin C Y . ROUGE: a package for automatic evaluation of summaries [C]// Proceedings of Workshop on Text Summarization Branches Out. Stroudsburg, PA: ACL, 2004: 74-81.

[34]

Li C Y, Liang X D, Hu Z T, et al. Hybrid retrieval—generation reinforced agent for medical image report generation [C]// Proceedings of the 32nd International Conference on Neural Information Processing Systems. Red Hook, NY: Curran Associates Inc., 2018: 1537-1547.

[35]

Jing B Y, Wang Z Y, Xing E . Show, describe and conclude: on exploiting the structure information of chest X—ray reports [C]// Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA: ACL, 2019: 6570-6580.

[36]

Li C Y, Liang X D, Hu Z T, et al. Knowledge—driven encode, retrieve, paraphrase for medical image report generation [J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2019, 33(1): 6666-6673.

[37]

Chen Z H, Song Y, Chang T H, et al. Generating radiology reports via memory—driven transformer [C]// Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, PA: ACL, 2020: 1439-1449.

[38]

Liu F L, Yin C C, Wu X, et al. Contrastive attention for automatic chest X—ray report generation [C]// Findings of the Association for Computational Linguistics: ACL—IJCNLP 2021. Stroudsburg, PA: ACL, 2021: 269-280.

基金资助

首都卫生发展科研专项资助项目(首发2024-2-2054)

AI Summary AI Mindmap
PDF (3294KB)

0

访问

0

被引

详细

导航
相关文章

AI思维导图

/