MFDU-Net:基于改进 U 型网络的肾脏超声图像分割

刘育铭 ,  代煜 ,  周璐 ,  张建勋

工程科学学报 ›› 2026, Vol. 48 ›› Issue (8) : 1830 -1841.

PDF
工程科学学报 ›› 2026, Vol. 48 ›› Issue (8) : 1830 -1841. DOI: 10.13374/j.issn2095-9389.2025.11.24.002
信息工程·控制科学与工程

MFDU-Net:基于改进 U 型网络的肾脏超声图像分割

作者信息 +

MFDU-Net: kidney ultrasound image segmentation based on improved U-shaped network

Author information +
文章历史 +
PDF

摘要

超声图像分割在确保准确诊断和制订有效治疗方案方面起着至关重要的作用,目前已经开发了许多深度学习方法来分割超声图像中的器官和病变. 然而,模糊的组织边界和大量的散斑噪声等问题限制了现有模型在超声图像上的分割精度. 为了解决这些问题,本文在 U-Net 模型基础上进行改进,提出了一种基于多尺度特征提取和频率注意力去噪的网络模型 MFDU-Net. 首先,在图像输入端采用了多分辨率输入模块,使模型能够为不同编码层提供输入信息. 其次,将原始 U-Net 的卷积编码层替换为多尺度分离卷积模块,以捕捉不同尺度的图像特征信息. 最后,在跳跃连接处引入频率去噪注意力模块,克服大多数医学图像分割网络中常见的空间域特征学习的局限性,减少无关信息和噪声对模型分割性能的影响. 对所提出的 MFDU-Net 在两个自建肾脏超声数据集及 DDTI 和 ISIC2018 两个公开数据集上进行了对比评估,与基线 U-Net 方法相比,MFDU-Net 在 Dice 指标上分别提高了 5.47、9.66、4.83 和 3.59 个百分点,HD95 距离指标分别缩减了 36.59、5.76、48.25 和 25.75,取得了良好的分割效果. 实验结果表明,与现有的先进医学图像分割模型相比,MFDU-Net 在分割器官和病灶方面具有更优越的性能.

Abstract

Ultrasound image segmentation is critical for accurate diagnosis and effective treatment planning. Although numerous deep learning methods have been developed for organ and lesion segmentation in ultrasound images, their performance remains limited by indistinct anatomical boundaries and high speckle noise. To solve these problems, this study improves the U-Net model and proposes a network model, MFDU-Net, based on multiscale feature extraction and frequency attention denoising. First, a multi-resolution input (MI) module is introduced at the network input end to select pooling windows of different sizes for size reduction from the feature maps extracted after one convolution of the input image. The feature maps are then overlaid with the corresponding-size encoding layers by channels for feature fusion, and the model provides input information for different encoding layers for subsequent encoding operations. Second, the convolutional encoding layer of the original U-Net is replaced with a multi-scale separation convolution (MSC) module to capture the image feature information at different scales. Unlike existing multiscale methods that use different sizes of convolution kernels in parallel, the proposed framework employs multiple feature extraction branches with the same 3×3 convolution kernel. By changing the number of convolutions, the same effect as changing the size of the convolution kernel is achieved, enabling multiscale feature extraction and allowing the network to adapt to targets of different sizes and positions, thus enhancing its expressive power and robustness. Finally, considering that the U-Net skip connection does not consider the direct transmission of noise and irrelevant information when fusing the encoder and decoder features, a frequency-denoising attention module is introduced in the skip connection. Drawing on the ideas of the CBAM module, frequency–space attention and frequency–channel attention modules are designed to remove noise and weight feature information, respectively. This overcomes the limitations of spatial domain feature learning commonly found in most medical image segmentation networks and reduces the impact of irrelevant information and noise on the model segmentation performance. U-Net, DeepLabv3+, Att U-Net, ACC-Unet, and FCRNet were used as comparative methods to evaluate the proposed MFDU-Net on two self-built kidney ultrasound datasets and two publicly available datasets: DDTI and ISIC2018. Considering the small number of collected images and the requirement of segmentation models for the training set size, data augmentation was performed on all the datasets used in this study. Six quantitative evaluation indicators were selected: Jaccard coefficient, recall rate, accuracy rate, Dice coefficient, HD95, and ASSD. Compared with the baseline U-Net, MFDU-Net improves the dice index by 5.47, 9.66, 4.83, and 3.59 percentage points, respectively, and the HD95 distance index decreased by 36.59, 5.76, 48.25 and 25.75, respectively, achieving good segmentation results. To verify the effectiveness of the improved modules designed in this study, ablation experiments were conducted using a kidney ultrasound dataset. The experimental results show that, compared with existing advanced medical image segmentation models, MFDU-Net demonstrates superior performance in segmenting organs and lesions.

关键词

超声图像分割 / 多尺度融合 / 注意力机制 / 深度学习 / 图像处理

Key words

ultrasound image segmentation / multi-scale fusion / attention mechanism / deep learning / image processing

引用本文

引用格式 ▾
刘育铭,代煜,周璐,张建勋. MFDU-Net:基于改进 U 型网络的肾脏超声图像分割[J]. 工程科学学报, 2026, 48(8): 1830-1841 DOI:10.13374/j.issn2095-9389.2025.11.24.002

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

Gupta D, Anand R S. A hybrid edge—based segmentation approach for ultrasound medical images[J]. Biomed Signal Process Control, 2017, 31: 116

[2]

Hermawati F A, Tjandrasa H, Suciati N. Phase—based thresholding schemes for segmentation of fetal thigh cross—sectional region in ultrasound images[J]. J King Saud Univ Comput Inf Sci, 2022, 34(7): 4448

[3]

Jardim S, António J, Mora C. Image thresholding approaches for medical image segmentation — short literature review[J]. Procedia Comput Sci, 2023, 219: 1485

[4]

Shao L P, Zhou Z B, Wu H M, et al. Modeling of hidden Markov in ultrasound image—assisted diagnosis[J]. J Healthcare Eng, 2021, 2021: 5597591

[5]

Shu X, Yang Y Y, Liu J, et al. ALVLS: Adaptive local variances—Based levelset framework for medical images segmentation[J]. Pattern Recognit, 2023, 136: 109257

[6]

He A L, Wang K, Li T, et al. Progressive multiscale consistent network for multiclass fundus lesion segmentation[J]. IEEE Trans Med Imaging, 2022, 41(11): 3146

[7]

Zhang S S, Liu R X, Shan K, et al. Research on cardiac image segmentation algorithm based on improved level set modeling[J]. Chin J Eng, 2025, 47(7): 1536

[8]

(张帅帅, 刘瑞霞, 单珂, . 基于改进水平集模型的心脏图像分割算法[J]. 工程科学学报, 2025, 47(7): 1536)

[9]

Wang R S, Lei T, Cui R X, et al. Medical image segmentation using deep learning: A survey[J]. IET Image Process, 2022, 16(5): 1243

[10]

Long J, Shelhamer E, Darrell T. Fully convolutional networks for semantic segmentation[C]// 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015: 3431

[11]

Ronneberger O, Fischer P, Brox T. U—Net: Convolutional networks for biomedical image segmentation[C]// Medical Image Computing and Computer—Assisted Intervention—MICCAI 2015, 2015: 234

[12]

Chen G P, Li L, Zhang J X, et al. Rethinking the unpretentious U—Net for medical ultrasound image segmentation[J]. Pattern Recognit, 2023, 142: 109728

[13]

Xiao X, Lian S, Luo Z M, et al. Weighted res—UNet for high—quality retina vessel segmentation[C]// 2018 9th International Conference on Information Technology in Medicine and Education (ITME), 2018: 327

[14]

Liu X, Gao P, Yu T, et al. CSWin—UNet: Transformer UNet with cross—shaped windows for medical image segmentation[J]. Inf Fusion, 2025, 113: 102634

[15]

Oghli M G, Bagheri S M, Shabanzadeh A, et al. Fully automated kidney image biomarker prediction in ultrasound scans using Fast—Unet++[J]. Sci Rep, 2024, 14(1): 4782

[16]

Zhou Z W, Rahman Siddiquee M M, Tajbakhsh N, et al. UNet++: A nested U—Net architecture for medical image segmentation[C]// Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support, 2018: 3

[17]

Huang H M, Lin L F, Tong R F, et al. UNet 3+: A full—scale connected UNet for medical image segmentation[C]// ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020: 1055

[18]

Chen S H, Wu Y L, Pan C Y, et al. Renal ultrasound image segmentation method based on channel attention and GL—UNet11[J]. J Radiat Res Appl Sci, 2023, 16(3): 100631

[19]

Oktay O, Schlemper J, Le Folgoc L, et al. Attention U—Net: Learning where to look for the pancreas [PP/OL]. arXiv (2018—04—11)[2025—11—24]. https://arxiv.org/abs/1804.03999

[20]

Li C, Tan Y S, Chen W, et al. Attention Unet++: A nested attention—aware U—Net for liver CT image segmentation[C]// 2020 IEEE International Conference on Image Processing (ICIP), 2020: 345

[21]

Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16×16 words: Transformers for image recognition at scale [PP/OL]. arXiv (2020—10—22)[2025—11—24]. https://arxiv.org/abs/2010.11929

[22]

Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[C]// 31st Conference on Neural Information Processing Systems (NIPS 2017), 2017

[23]

Yang B C, Wang J Y, Jin H B. DS—TransFusion: Automatic retinal vessel segmentation based on an improved Swin Transformer[J]. Chin J Eng, 2024, 46(10): 1889

[24]

(杨本臣, 王建宇, 金海波. DS—TransFusion:基于改进 Swin Transformer 的视网膜血管自动分割[J]. 工程科学学报, 2024, 46(10): 1889)

[25]

Liang L M, He A J, Yang Y, et al. Colorectal polyp segmentation method based on the Swin Transformer and graph reasoning[J]. Chin J Eng, 2024, 46(5): 897

[26]

(梁礼明, 何安军, 阳渊, . 基于 Swin Transformer 和图形推理的结直肠息肉分割方法[J]. 工程科学学报, 2024, 46(5): 897)

[27]

Chen J N, Mei J R, Li X H, et al. TransUNet: Rethinking the U—Net architecture design for medical image segmentation through the lens of transformers[J]. Med Image Anal, 2024, 97: 103280

[28]

Ozcan A, Tosun Ö, Donmez E, et al. Enhanced—TransUNet for ultrasound segmentation of thyroid nodules[J]. Biomed Signal Process Control, 2024, 95: 106472

[29]

Fu H Z, Cheng J, Xu Y W, et al. Joint optic disc and cup segmentation based on multi—label deep network and polar transformation[J]. IEEE Trans Med Imaging, 2018, 37(7): 1597

[30]

Gao S H, Cheng M M, Zhao K, et al. Res2Net: A new multi—scale backbone architecture[J]. IEEE Trans Pattern Anal Mach Intell, 2021, 43(2): 652

[31]

He K M, Zhang X Y, Ren S Q, et al. Deep residual learning for image recognition[C]// 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016: 770

[32]

Woo S, Woo S, Park J, et al. CBAM: Convolutional block attention module[C]// Computer Vision—ECCV 2018, 2018: 3

[33]

Tang R, Zhao H J, Tong Y, et al. A frequency attention—embedded network for polyp segmentation[J]. Sci Rep, 2025, 15: 4961

[34]

Patro B N, Namboodiri V P, Agneeswaran V S. SpectFormer: Frequency and attention is what you need in a vision transformer[C]// 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2025: 9543

[35]

Ahmed N, Natarajan T, Rao K R. Discrete cosine transform[J]. IEEE Trans Comput, 1974, C—23(1): 90

[36]

Qin Z Q, Zhang P Y, Wu F, et al. FcaNet: Frequency channel attention networks[C]// 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2021: 763

[37]

Chen L C, Zhu Y K, Papandreou G, et al. Encoder—decoder with atrous separable convolution for semantic image segmentation[C]// Computer Vision—ECCV 2018, 2018: 833

[38]

Ibtehaz N, Kihara D. ACC—UNet: A completely convolutional UNet model for the 2020s[C]// Medical Image Computing and Computer Assisted Intervention—MICCAI 2023, 2023: 692

[39]

He A L, Li T, Wu Y L, et al. FRCNet: Frequency and region consistency for semi—supervised medical image segmentation[C]// Medical Image Computing and Computer Assisted Intervention—MICCAI 2024, 2024: 305

基金资助

国家自然科学基金资助项目(62173190)

AI Summary AI Mindmap
PDF

0

访问

0

被引

详细

导航
相关文章

AI思维导图

/