融合双分支编码与扩散先验的动态舌体配准研究

汤致文, 余瀚, 姜成惠, 夏春雨, 李平

小型微型计算机系统 ›› 2026, Vol. 47 ›› Issue (9) : 2211 -2219.

小型微型计算机系统 ›› 2026, Vol. 47 ›› Issue (9) : 2211 -2219. DOI: 10.20009/j.cnki.21-1106/TP.2025-0409
计算机图形与图像

融合双分支编码与扩散先验的动态舌体配准研究

    汤致文1,3, 余瀚1,3, 姜成惠2,3, 夏春雨2, 李平1
作者信息 +

Integrating Dual-branch Encoding and Diffusion Prior for Dynamic Tongue Registration

    TANG Zhiwen1,3, YU Han1,3, JIANG Chenghui2,3, XIA Chunyu2, LI Ping1
Author information +
文章历史 +

摘要

医学图像配准是表征舌体运动模式、评估术后功能恢复情况的关键技术.目前对舌体配准研究不足,且舌体在运动中的非规则性及局部异质性形变特征的建模能力面临挑战.本文借助舌体动态时序核磁影像,提出了一种融合可变形卷积与窗口注意力机制的双分支编码器,以提取舌体轮廓的局部几何形变特征与全局语义上下文信息,并引入预训练潜在扩散模型提取的多尺度层级先验特征进行增强,从而实现对舌体复杂运动变形模式的表征.实验表明,本文网络在冠状位、矢状位舌体核磁数据集上取得了dice=0.9795、0.9856优于现有模型的指标表现;同时,在2心室、4心室心脏超声数据集上取得了dice=0.8805、0.8964最优表现,验证了模型的泛化能力.

Abstract

Medical image registration is a key technique for characterizing tongue motion patterns and assessing postoperative functional recovery.However,current research on tongue registration remains insufficient,and the ability to model the non-regular and locally heterogeneous deformation features of the tongue during motion poses a significant challenge.Leveraging dynamic magnetic resonance imaging(Cine-MRI) of the tongue,this paper proposes a dual-branch encoder that integrates deformable convolution and windowed attention mechanisms.This architecture extracts both local geometric deformation features and global semantic context information of the tongue contour in parallel.Furthermore,multi-scale hierarchical prior features derived from a pre-trained latent diffusion model are incorporated to enhance the representation of the tongue′s complex motion deformation patterns.Experimental results demonstrate that the proposed network achieves Dice scores of 0.9795 and 0.9856 on coronal and sagittal tongue MRI datasets,respectively,outperforming existing models.It also achieves state-of-the-art Dice scores of 0.8805 and 0.8964 on the 2-chamberand 4-chambercardiac ultrasound datasets datasets,confirming its strong generalization capability.

关键词

医学图像配准 / 深度学习 / 自注意力 / 可变卷积 / 扩散模型

Key words

medical image registration / deep learning / self-attention / deformable convolution / diffusion model

引用本文

引用格式 ▾
汤致文, 余瀚, 姜成惠, 夏春雨, 李平. 融合双分支编码与扩散先验的动态舌体配准研究[J]. 小型微型计算机系统, 2026, 47(9): 2211-2219 DOI:10.20009/j.cnki.21-1106/TP.2025-0409

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1] Wu J,Chen H,Liu Y,et al.The global,regional,and national burden of oral cancer,1990~2021:a systematic analysis for the global burden of disease study 2021[J].Journal of Cancer Research and Clinical Oncology,2025,151(2):53.
[2] Kreeft A M,Molen L,Hilgers F J,et al.Speech and swallowing after surgical treatment of advanced oral and oropharyngeal carcinoma:a systematic review of the literature[J].European Archives of Oto-Rhino-Laryngology,2009,266(11):1687-1698.
[3] Lin P C,Wang W C,Kao Y H,et al.Long-term effects of an oral rehabilitation programme on the oral function of male patients with or without tongue cancer[J].Journal of Oral Rehabilitation,2025,doi:10.1111/joor.14043.
[4] Chauvel M,Tessier C,Venkatasamy A,et al.Cine-MRI of deglutition:asystematic review[J].Dysphagia,2025,40(4):700-710.
[5] XIAO L J,ZHANG L L,LIAO M X,et al.Reliability and validity of the Chinese version of the dynamic imaging grade of swallowing toxicity scale[J].Chinese Journal of RehabilitationMedicine,2020,35(3):6,doi:10.3969/j.issn.1001-1242.2020.03.004.
[6] MO X Y,YANG F,YIN M X,et al.Deep learning methods for medical image registration:a survey[J].Journal of Chinese Computer Systems,2021,42(8):1706-1714.
[7] Wang W,Dai J,Chen Z,et al.Internimage:exploring large-scale vision foundation models with deformable convolutions[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2023:14408-14419.
[8] Liu Z,Lin Y,Cao Y,et al.Swin transformer:hierarchical vision transformer using shifted windows[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision,2021:10012-10022.
[9] Vercauteren T,Pennec X,Perchant A,et al.Non-parametric diffeomorphic image registration with the demons algorithm[C]//International Conference on Medical Image Computing and Computer-assisted Intervention,2007:319-326.
[10] Beg M F,Miller M I,Trouvé A,et al.Computing large deformation metric mappings via geodesic flows of diffeomorphisms[J].International Journal of Computer Vision,2005,61(2):139-157.
[11] Shackleford J A,Kandasamy N,Sharp G C.On developing B-spline registration algorithms for multi-core processors[J].Physics in Medicine and Biology,2010,55(21):6329-6351.
[12] Balakrishnan G,Zhao A,Sabuncu M R,et al.Voxelmorph:a learning framework for deformable medical image registration[J].IEEE Transactions on Medical Imaging,2019,38(8):1788-1800.
[13] Li R,Figueredo G,Auer D,et al.Mrregnet:multi-resolution mask guided convolutional neural network for medical image registration with large deformations[C]//IEEE International Symposium on Biomedical Imaging(ISBI),2024:1-5.
[14] Han K,Xiao A,Wu E,et al.Transformer in transformer[C]//Advances in Neural Information Processing Systems,2021:15908-15919.
[15] Chen J,Frey E C,He Y,et al.Transmorph:transformer for unsupervised medical image registration[J].Medical Image Analysis,2022,82:102615,doi:10.1016/j.media.2022.102615.
[16] Kim B,Han I,Ye J C.Diffusemorph:unsupervised deformable image registration using diffusion model[C]//European Conference on Computer Vision,2022:347-364.
[17] Wu J,Gong K.LDM-Morph:latent diffusion model guided deformable image registration[J].arXiv preprint arXiv:2411.15426,2024.
[18] Craig T,Battista J,Dyk J V.Limitations of a convolution method for modeling geometric uncertainties in radiation therapy.I.The effect of shift invariance[J].Medical Physics,2003,30(8):2001-2011.
[19] Hou Q,Zhou D,Feng J.Coordinate attention for efficient mobile network design[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2021:13713-13722.
[20] WANG Z,NI J H,ZHANG F L.Global-aware fusion of binary multi-scale features for speech emotion recognition[J].Journal of Computer Applications,2025,44(10):1-9.
[21] Rombach R,Blattmann A,Lorenz D,et al.High-resolution image synthesis with latent diffusion models[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2022:10684-10695.
[22] YANG C R,SHEN H Y,CHE H,et al.Seismic surface wave separation based on shifted window self-attention network and transfer learning[J].Journal of Xi′an Shiyou University(Natural Science Edition),2024,39(6):39-50.
[23] Zhu X,Hu H,Lin S,et al.Deformable convnets v2:more deformable,better results[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2019:9308-9316.
附中文参考文献:
[5] 肖灵君,张路路,廖美新,等.中文版吞咽障碍造影评分量表的信度和效度研究[J].中国康复医学杂志,2020,35(3):6,doi:10.3969/j.issn.1001-1242.2020.03.004.
[6] 莫晓盈,杨 锋,尹梦晓,等.医学图像配准的深度学习方法综述[J].小型微型计算机系统,2021,42(8):1706-1714.
[20] 王 政,倪佳慧,张凡龙.全局感知融合二值多尺度特征的语音情感识别[J].计算机应用,2025,44(10):1-9.
[22] 杨晨睿,沈鸿雁,车 晗,等.基于位移窗口自注意力网络和迁移学习的地震面波分离[J].西安石油大学学报(自然科学版),2024,39(6):39-50.

基金资助

国家自然科学基金项目(12371440)资助;南京邮电大学自然科学基金项目(NY224142)资助;口腔疾病研究与防治国家级重点实验室培育建设点开放课题基金项目(JSKLOD-KF-2506)资助.

AI Summary AI Mindmap

0

访问

0

被引

详细

导航
相关文章

AI思维导图

/