基于跨模态生成与交叉注意力的乳腺超声诊断
傅力嘉 , 林艳萍 , 李娜 , 贾超 , 李刚 , 杜联芳 , 李凡
中国医学物理学杂志 ›› 2026, Vol. 43 ›› Issue (7) : 889 -897.
基于跨模态生成与交叉注意力的乳腺超声诊断
Breast ultrasound diagnosis based on cross-modal generation and cross-attention
目的:针对常规B超在乳腺癌筛查中因读图者间一致性不佳、微小病灶恶性征象不典型导致的准确性受限问题,以及超声造影因成本高、操作复杂而难以普及的临床矛盾,提出一种“分割-生成-分类”三阶段级联诊断框架,旨在仅利用常规B超图像实现高精度的乳腺肿瘤良恶性诊断。方法:首先,采用CNN-Transformer架构的分割网络从B超图像中提取肿瘤区域。接着,构建基于条件生成对抗网络的跨模态生成模型,利用预训练的ResNet-34编码器从分割后的B超中生成对应的超声造影图像。最后,设计以Inception-v3为骨干的双分支交叉注意力网络,通过动态特征融合实现对真实B超与生成造影信息的深度交互,完成对肿瘤良恶性的精确诊断。结果:在包含677例患者的乳腺超声数据集上进行验证,该框架的受试者工作特征曲线下面积(AUC)达到0.930,显著优于多个单模态B超基线模型(最佳AUC为0.868)。消融实验证实病灶分割、跨模态生成以及交叉注意力特征融合模块对提升诊断性能的必要贡献。结论:本研究提出的框架验证了通过深度学习方法生成并利用多模态超声信息进行高精度诊断的可行性与有效性,能够以低成本、高效率的方式提供多模态诊断增益,有望为乳腺癌临床筛查提供有力的辅助决策支持,降低乳腺癌临床筛查的漏诊与误诊风险。
Objective To address the limited accuracy of conventional B-mode ultrasound (US) in breast cancer screening, resulting from poor inter-reader agreement and atypical malignant signs of tiny lesions, and resolve the clinical dilemma where contrast-enhanced ultrasound (CEUS) is hard to popularize due to its high cost and operational complexity, this study proposes a three-stage cascaded diagnostic framework of "segmentation-generation-classification" to achieve high-precision diagnosis of benign and malignant breast tumors relying solely on conventional US images. Methods A CNN-Transformer framework-based segmentation network was first employed to extract tumor regions from US images. Subsequently, a cross-modal generation model based on a conditional generative adversarial network was constructed, utilizing a pre-trained ResNet-34 encoder to synthesize corresponding CEUS images from the segmented US images. Finally, a dual-branch cross-attention network with Inception-v3 as the backbone was designed to achieve deep interaction between the real US and the generated CEUS information through dynamic feature fusion, thereby accomplishing accurate benign and malignant tumor classification. Results Validated on a breast ultrasound dataset including 677 patients, the proposed framework achieved an area under the receiver operating characteristic curve (AUC) of 0.930, significantly outperforming several single-modal US baseline models which reached a maximum AUC of 0.868. Ablation experiments verified the essential contributions of lesion segmentation, cross-modal generation, and cross-attention feature fusion modules in enhancing diagnostic performance. Conclusion This study validates the feasibility and effectiveness of generating and utilizing multimodal ultrasound information with deep learning for high-precision diagnosis. The proposed framework provides multimodal diagnostic benefits in a cost-effective and highly efficient way, holding great promise to provide robust auxiliary decision support for clinical breast cancer screening and reduce the risks of missed diagnoses and misdiagnoses.
| [1] |
|
| [2] |
|
| [3] |
|
| [4] |
|
| [5] |
|
| [6] |
|
| [7] |
|
| [8] |
|
| [9] |
|
| [10] |
|
| [11] |
|
| [12] |
|
| [13] |
孙雨, 杨琛 . 基于自动乳腺超声诊断系统的乳腺肿瘤人工智能诊断研究进展[J]. 中国医学影像学杂志, 2024, 32(11): 1176-1181. |
| [14] |
|
| [15] |
|
| [16] |
王一凡, 刘静, 马金刚, |
| [17] |
|
| [18] |
|
| [19] |
|
| [20] |
柴梦婷, 朱远平 . 生成式对抗网络研究与应用进展[J]. 计算机工程, 2019, 45(9): 222-234. |
| [21] |
|
| [22] |
|
| [23] |
|
| [24] |
|
| [25] |
|
| [26] |
|
| [27] |
|
| [28] |
|
| [29] |
何俊, 张彩庆, 李小珍, |
| [30] |
|
| [31] |
|
| [32] |
|
| [33] |
|
| [34] |
|
| [35] |
|
国家自然科学基金(52175020)
上海市自然科学基金(21ZR1451400)
/
| 〈 |
|
〉 |