先验匹配激活和注意力特征融合的小样本语义分割方法

王熠聪 ,  黄荣 ,  蒋学芹 ,  周树波

东华大学学报(自然科学版) ›› 2026, Vol. 52 ›› Issue (2) : 164 -171.

PDF (9358KB)
东华大学学报(自然科学版) ›› 2026, Vol. 52 ›› Issue (2) : 164 -171. DOI: 10.19886/j.cnki.dhdz.2025.0026
信息与智能科学及纺织智能制造

先验匹配激活和注意力特征融合的小样本语义分割方法

作者信息 +

Prior matching activation and attention-based feature fusion for few-shot semantic segmentation

Author information +
文章历史 +
PDF (9582K)

摘要

小样本语义分割旨在利用少量带标注的支持图像,分割出查询图像中的目标。现有方法采用特征平均提取用于分割查询图像的支持原型,易导致语义信息丢失,且直接排除背景特征,特征利用率低。为缓解上述问题,提出一种先验匹配激活和注意力特征融合的小样本语义分割方法。将支持图像的前景特征作为先验知识,通过免训练的先验匹配激活实现查询图像前景的粗略定位,形成先验掩码。在此引导下,基于交叉注意力机制以可学习的方式融合支持图像的前景和背景特征。这种特征融合方式兼顾前景与背景特征,取代了不可学习的掩码平均操作,特征利用率高,有利于缓解目标语义的丢失问题。试验结果表明:在 PASCAL-5i 和 COCO-20i 数据集上,所提方法优于现有小样本语义分割方法,1-shot的平均交并比(mIoU)分别提高了19.00% 和64.47%;5-shot的平均交并比 mIoU 分别提高16.00%和44.99%。

Abstract

Few-shot semantic segmentation (FSS) aims to segment target object in a query image using only a few labeled support images. Existing methods adopt feature averaging to extract support prototypes for segmenting the query image, which leads to the loss of target semantics and directly excludes background features, resulting in inefficient feature exploitation. To address these issues, a few-shot semantic segmentation method based on prior matching activation and attention-based feature fusion is proposed. The foreground features of the support image are treated as prior knowledge, and prior matching activation is performed in a training-free manner to achieve a coarse localization of the foreground of the query image, generating a prior mask. Guided by this prior mask, the foreground and background features of the support image are fused using a learnable cross-attention mechanism. This feature fusion approach considers both foreground and background features, replacing the non-learnable mask averaging operation, improving efficiency of feature exploitation, and alleviating the issue of target semantic loss. Experimental results demonstrate that the proposed method outperforms existing FSS methods on the PASCAL-5i and COCO-20i datasets. The proposed method achieves an average mean Intersection over Union (mIoU) improvement of 19.00% and 64.47% for 1-shot, and 16.00% and 44.99% for 5-shot, respectively.

关键词

小样本语义分割 / 先验匹配激活 / 注意力特征融合 / 掩码平均 / 原型学习

Key words

few-shot semantic segmentation / prior matching activation / attention-based feature fusion / mask averaging / prototype learning

引用本文

引用格式 ▾
王熠聪,黄荣,蒋学芹,周树波. 先验匹配激活和注意力特征融合的小样本语义分割方法[J]. 东华大学学报(自然科学版), 2026, 52(2): 164-171 DOI:10.19886/j.cnki.dhdz.2025.0026

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

王晓, 张翔宇, 周锐, . 基于平行测试的认知自动驾驶智能架构研究[J]. 自动化学报, 2024, 50(2): 356-371.

[2]

WANG X, ZHANG X Y, ZHOU R, et al. An intelligent architecture for cognitive autonomous driving based on parallel testing[J]. Acta Automatica Sinica, 2024, 50(2): 356-371.

[3]

CHENG K, SUN D Y, CHEN C, et al. Intelligent gear decision method for automatic vehicles based on data mining under uphill conditions[J]. IEEE Transactions on Intelligent Transportation Systems, 2022, 23(12): 24235-24247.

[4]

WANG Z G, ZHAN J, DUAN C G, et al. A review of vehicle detection techniques for intelligent vehicles[J]. IEEE Transactions on Neural Networks and Learning Systems, 2023, 34(8): 3811-3831.

[5]

田娟秀, 刘国才, 谷珊珊, . 医学图像分析深度学习方法研究与挑战[J]. 自动化学报, 2018, 44(3): 401-424.

[6]

TIAN J X, LIU G C, GU S S, et al. Deep learning in medical image analysis and its challenges[J]. Acta Automatica Sinica, 2018, 44(3): 401-424.

[7]

SHEN D G, WU G R, SUK H I. Deep learning in medical image analysis[J]. Annual Review of Biomedical Engineering, 2017, 19: 221-248.

[8]

WANG P J, BAYRAM B, SERTEL E. A comprehensive review on deep learning based remote sensing image super-resolution methods[J]. Earth-Science Reviews, 2022, 232: 104110.

[9]

WANG Y, ALBRECHT C M, ALI BRAHAM N A, et al. Self-supervised learning in remote sensing: a review[J]. IEEE Geoscience and Remote Sensing Magazine, 2022, 10(4): 213-247.

[10]

CHEN L C, PAPANDREOU G, KOKKINOS I, et al. DeepLab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, 40(4): 834-848.

[11]

LIN G S, LIU F Y, MILAN A, et al. RefineNet: multi-path refinement networks for dense prediction[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 42(5): 1228-1242.

[12]

SHELHAMER E, LONG J, DARRELL T. Fully convolutional networks for semantic segmentation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(4): 640-651.

[13]

ZHAO H S, ZHANG Y, LIU S, et al. PSANet: point-wise spatial attention network for scene parsing[C]// Computer Vision-ECCV 2018. Cham: Springer, 2018: 270-286.

[14]

ZHANG X L, WEI Y C, YANG Y, et al. SG-one: similarity guidance network for one-shot semantic segmentation[J]. IEEE Transactions on Cybernetics, 2020, 50(9): 3855-3865.

[15]

SHABAN A, BANSAL S, LIU Z, et al. One-shot learning for semantic segmentation[C]// Proceedings of the British Machine Vision Conference 2017. London, UK. British Machine Vision Association, 2017.

[16]

DONG N Q, XING E P. Few-shot semantic segmentation with prototype learning[C]// British Machine Vision Conference, 2018.

[17]

RAKELLY K, SHELHAMER E, DARRELL T, et al. Conditional networks for few-shot semantic segmentation[C]// Proceedings of the International Conference on Learning Representations. Vancouver, Canada: 2018.

[18]

NGUYEN K, TODOROVIC S. Feature weighting and boosting for few-shot segmentation[C]// 2019 IEEE/CVF International Conference on Computer Vision (ICCV). Seoul, Korea. IEEE, 2019: 622-631.

[19]

LIU Y F, ZHANG X Y, ZHANG S Y, et al. Part-aware prototype network for few-shot semantic segmentation[C]// Computer Vision-ECCV 2020. Cham: Springer, 2020: 142-158.

[20]

WANG K X, LIEW J H, ZOU Y T, et al. PANet: few-shot image semantic segmentation with prototype alignment[C]// 2019 IEEE/CVF International Conference on Computer Vision (ICCV). Seoul, Korea. IEEE, 2019: 9196-9205.

[21]

TIAN Z T, ZHAO H S, SHU M, et al. Prior guided feature enrichment network for few-shot segmentation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(2): 1050-1065.

[22]

GAO G Y, FANG Z Y, HAN C, et al. DRNet: double recalibration network for few-shot semantic segmentation[J]. IEEE Transactions on Image Processing, 2022, 31: 6733-6746.

[23]

FAN Q, PEI W J, TAI Y W, et al. Self-support few-shot semantic segmentation[C]// Computer Vision-ECCV 2022. Cham: Springer, 2022: 701-719.

[24]

CONG R M, XIONG H, CHEN J P, et al. Query-guided prototype evolution network for few-shot segmentation[J]. IEEE Transactions on Multimedia, 2024, 26: 6501-6512.

[25]

YANG B Y, WAN F, LIU C, et al. Part-based semantic transform for few-shot semantic segmentation[J]. IEEE Transactions on Neural Networks and Learning Systems, 2022, 33(12): 7141-7152.

[26]

ZHANG X L, WEI Y C, LI Z, et al. Rich embedding features for one-shot semantic segmentation[J]. IEEE Transactions on Neural Networks and Learning Systems, 2022, 33(11): 6484-6493.

[27]

ZHANG C, LIN G S, LIU F Y, et al. CANet: class-agnostic segmentation networks with iterative refinement and attentive few-shot learning[C]// 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Long Beach, CA, USA. IEEE, 2020: 5212-5221.

[28]

LI G, JAMPANI V, SEVILLA-LARA L, et al. Adaptive prototype learning and allocation for few-shot segmentation[C]// 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, TN, USA. IEEE, 2021: 8330-8339.

[29]

LIU Y W, LIU N, CAO Q L, et al. Learning non-target knowledge for few-shot semantic segmentation[C]// 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans, LA, USA. IEEE, 2022: 11563-11572.

[30]

MIN J H, KANG D, CHO M. Hypercorrelation squeeze for few-shot segmenation[C]// 2021 IEEE/CVF International Conference on Computer Vision (ICCV). 2021.

[31]

DENG J, DONG W, SOCHER R, et al. ImageNet: a large-scale hierarchical image database[C]// 2009 IEEE Conference on Computer Vision and Pattern Recognition. Miami, FL, USA. IEEE, 2009: 248-255.

[32]

VASWANI A. Attention is all you need[C]// Proceedings of the Advances in Neural Information Processing Systems. Long Beach, CA, USA: 2017: 6000-6010.

[33]

EVERINGHAM M, VAN GOOL L, WILLIAMS C K I, et al. The pascal visual object classes (VOC) challenge[J]. International Journal of Computer Vision, 2010, 88(2): 303-338.

[34]

HARIHARAN B, ARBELÁEZ P, BOURDEV L, et al. Semantic contours from inverse detectors[C]// 2011 International Conference on Computer Vision. Barcelona, Spain. IEEE, 2012: 991-998.

[35]

LIN T Y, MAIRE M, BELONGIE S, et al. Microsoft COCO: common objects in context[C]// Computer Vision-ECCV 2014. Cham: Springer, 2014: 740-755.

基金资助

国家自然科学基金(62001099)

中央高校基本科研业务费专项资金(2232023D-30)

AI Summary AI Mindmap
PDF (9358KB)

58

访问

0

被引

详细

导航
相关文章

AI思维导图

/