联合全局建模与局部补偿的高分辨率遥感影像语义分割

张建昌 ,  徐伟铭 ,  余大文 ,  梁昌钰 ,  陈枭

海南大学学报(自然科学版中英文) ›› 2026, Vol. 44 ›› Issue (4) : 417 -429.

PDF (12429KB)
海南大学学报(自然科学版中英文) ›› 2026, Vol. 44 ›› Issue (4) : 417 -429. DOI: 10.65658/j.hndk.2025101501
时空智能

联合全局建模与局部补偿的高分辨率遥感影像语义分割

作者信息 +

Semantic segmentation for remote sensing images based on joint global feature modeling and local feature compensation

Author information +
文章历史 +
PDF (12726K)

摘要

为解决高分辨率遥感影像语义分割中普遍存在的局部细节刻画不足以及全局建模效率较低的问题,本文提出了一种结合卷积神经网络与状态空间模型的分割网络模型—残差曼巴 U 型网络模型。该模型采用轻量级的 ResNet18 作为编码器,并引入视觉状态空间模块构建解码器,实现高效的全局上下文建模。同时,为增强对细粒度语义信息的感知能力,设计了局部特征补偿模块。针对深浅层特征融合过程中易出现的语义偏差问题,进一步提出跨层级融合注意力模块,以实现空间与语义信息的有效协同。结果表明,该模型在 Vaihingen、Potsdam 和 LoveDA 三大遥感数据集上的平均交并比分别达到 83.78%、87.09% 和 52.85%,均优于现有的主流模型。该模型在维持较低计算复杂度的同时,显著提升了模型的特征表达能力与分割精度。研究结果为高分辨率遥感影像的高效智能解译提供了一种兼顾性能与效率的可行方案。

Abstract

To address the common challenges of insufficient local detail representation and low global modeling efficiency in high-resolution remote sensing image semantic segmentation, this paper proposes a segmentation network that integrates convolutional neural networks and state-space models, named RMUNet. The proposed model employs a lightweight ResNet18 as the encoder and introduces a visual state space block (VSSBlock) in the decoder to achieve efficient global context modeling. Meanwhile, a local feature compensation module (LFCM) is designed to enhance the perception of fine-grained semantic information. To mitigate the semantic bias that may arise during the fusion of shallow and deep features, a cross-level fusion attention module (CFAM) is further proposed to enable effective collaboration between spatial and semantic representations. Experimental results demonstrate that RMUNet achieves mean intersection-over-union (mIoU) scores of 83.78%, 87.09%, and 52.85% on the Vaihingen, Potsdam, and LoveDA datasets, respectively, outperforming existing mainstream methods. While maintaining low computational complexity, RMUNet significantly enhances feature representation and segmentation accuracy, providing an efficient and effective solution for high-resolution remote sensing image interpretation.

关键词

高分辨率遥感影像 / 语义分割 / Mamba / 卷积神经网络

Key words

high-resolution remote sensing images / semantic segmentation / Mamba / convolutional neural networks

引用本文

引用格式 ▾
张建昌,徐伟铭,余大文,梁昌钰,陈枭. 联合全局建模与局部补偿的高分辨率遥感影像语义分割[J]. 海南大学学报(自然科学版中英文), 2026, 44(4): 417-429 DOI:10.65658/j.hndk.2025101501

登录浏览全文

4963

注册一个新账户 忘记密码

作者贡献声明

张建昌负责论文的整体构思与实验方案设计,撰写论文初稿。徐伟铭提供研究经费。余大文负责全过程的论文指导。梁昌钰负责实验数据的分析与解释。陈枭负责实验数据的收集与处理。

AI使用声明

本文的英文题名、摘要和关键词采用腾讯元宝翻译生成,并进行了部分修改。

利益冲突声明

作者声明,不存在已知的可能影响本论文所报告工作的竞争性经济利益或个人关系。

伦理声明

本研究不涉及人类或动物的生物学研究。

数据可用性

支持论文结论所需的所有数据均为公开数据集。

参考文献

[1]

Liu Y CFan BWang L Fet al. Semantic labeling in very high resolution imagesvia a self—cascaded convolutional neural network [J]. ISPRS Journal of Photogrammetry and Remote Sensing2018145: 78-95.

[2]

Schumann G J PBrakenridge G RKettner A Jet al. Assisting flood disaster response with earth observation data and products:a critical assessment[J]. Remote Sensing201810(8): 1230.

[3]

Chen S W . SAR image speckle filtering with context covariance matrix formulation and similarity test[J]. IEEE Transactions on Image Processing202029: 6641-6654.

[4]

Zhao T JWang SOuyang C Jet al. Artificial intelligence for geoscience:progress,challenges,and perspectives[J]. The Innovation2024, 5(5): 100691.

[5]

Voulodimos ADoulamis NDoulamis Aet al. Deep learning for computer vision:a brief review[J]. Computational Intelligence and Neuroscience20182018(1): 7068349.

[6]

Vaswani AShazeer NParmar Net al. Attention is all you need[C]// Proceedings of the 31st International Conference on Neural Information Processing Systems. Long Beach:Curran Associates Inc. , 2017: 6000-6010.

[7]

Ronneberger OFischer PBrox T . U—Net:convolutional networks for biomedical image segmentation[C]// Proceedings of the 18th International Conference on Medical Image Computing and Computer—Assisted Intervention. Munich:Springer, 2015: 234-241.

[8]

Chen L CZhu Y KPapandreou Get al. Encoder—decoder with atrous separable convolution for semantic image segmentation[C]// Proceedings of the 15th European Conference on Computer Vision ( ECCV) . Munich: Springer, 2018: 833-851.

[9]

Zhao H SShi J PQi X Jet al. Pyramid scene parsing network[C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Honolulu:IEEE, 2017: 6230-6239.

[10]

Strudel RGarcia RLaptev Iet al. Segmenter:transformer for semantic segmentation[C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal:IEEE, 2021: 7242-7252.

[11]

Xie E ZWang W HYu Z Det al. SegFormer:simple and efficient design for semantic segmentation with transformers[C]// Advances in Neural Information Processing Systems 34 (NeurIPS 2021). New York: Curran Associates, 2021: 12077-12090.

[12]

Cao HWang Y YChen Jet al. Swin—Unet:Unet—like pure transformer for medical image segmentation[C]// Proceedings of the European Conference on Computer Vision. Tel Aviv:Springer, 2022: 205-218.

[13]

Liu ZLin Y TCao Yet al. Swin transformer:hierarchical vision transformer using shifted windows[C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal:IEEE, 2021: 9992-10002.

[14]

Chen J NLu Y YYu Q Het al. TransUNet:transformers make strong encoders for medical image segmentation[C]// Medical Image Computing and Computer Assisted Intervention—MICCAI 2021: 24th International Conference, Strasbourg, France, September 27—October 1, 2021, Proceedings, Part II. Cham: Springer International Publishing, 2021: 234-243.

[15]

Wang L BLi RDuan C Xet al. A novel transformer based semantic segmentation scheme for fine—resolution remote sensing images[J]. IEEE Geoscience and Remote Sensing Letters202219: 6506105.

[16]

Wang L BLi RWang D Zet al. Transformer meets convolution:a bilateral awareness network for semantic segmentation of very fine resolution urban scene images[J]. Remote Sensing202113(16): 3065.

[17]

Wang L BLi RZhang Cet al. UNetFormer:a UNet—like transformer for efficient semantic segmentation of remote sensing urban scene imagery[J]. ISPRS Journal of Photogrammetry and Remote Sensing2022190: 196-214.

[18]

Zhang H WZhu YWang Det al. A survey on visual Mamba[J]. Applied Sciences202414(13): 5683.

[19]

Gu ADao T . Mamba:linear—time sequence modeling with selective state spaces[C]// Proceedings of the 41st International Conference on Machine Learning (ICML 2024). Vienna, Austria: PMLR, 2024, 235: 16237-16280.

[20]

Gu AGoel K C . Efficiently modeling long sequences with structured state spaces[C]// Proceedings of the 10th International Conference on Learning Representations (ICLR 2022). Virtual Conference: OpenReview.net, 2022: uYLFoz1vlAC.

[21]

Liu YTian Y JZhao Y Zet al. VMamba:visual state space model[C]// Proceedings of the 38th International Conference on Neural Information Processing Systems. Vancouver:Curran Associates Inc. , 2024: 3273.

[22]

Zhao S JChen HZhang X Let al. RS—Mamba for large remote sensing image dense prediction[J]. IEEE Transactions on Geoscience and Remote Sensing202462: 5633314.

[23]

Zhu Q FCai Y ZFang Yet al. Samba:semantic segmentation of remotely sensed images with state space model[J]. Heliyon202410(19): e38495.

[24]

Zhu E ZChen ZWang D Ket al. UNetMamba:an efficient UNet—like Mamba for semantic segmentation of high—resolution remote sensing images[J]. IEEE Geoscience and Remote Sensing Letters202522: 6001205.

[25]

Ma X PZhang X KPun M O . RS3Mamba:visual state space model for remote sensing image semantic segmentation [J]. IEEE Geoscience and Remote Sensing Letters202421: 6011405.

[26]

Wang L BLi D XDong S Jet al. PyramidMamba:rethinking pyramid feature fusion with selective space state model for semantic segmentation of remote sensing imagery[J]. International Journal of Applied Earth Observation and Geoinformation2024144: 104884.

[27]

Liu M SDan JLu Z Qet al. CM—UNet:hybrid CNN—Mamba UNet for remote sensing image semantic segmentation[PP/OL]. arXiv (2024—05—16)[2025—07—01]. https://arxiv.org/abs/2405.10530.

[28]

Milletari FNavab NAhmadi S A . V—Net:fully convolutional neural networks for volumetric medical image segmentation[C]// Proceedings of the 4th International Conference on 3D Vision (3DV). Stanford:IEEE, 2016: 565-571.

[29]

Zhu ZXu M DBai Set al. Asymmetric non—local neural networks for semantic segmentation[C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Seoul:IEEE, 2019: 593-602.

AI Summary AI Mindmap
PDF (12429KB)

0

访问

0

被引

详细

导航
相关文章

AI思维导图

/