To solve the problems of insufficient understanding and fuzzy feature boundary segmentation of multi-target categories small-scale feature semantic information in aerial remote sensing images in complex background, this paper designs a segmentation model that integrates the features of the backbone network information and classifies and reconstructs the features to improve the segmentation effect. The model takes Swin-Transformer as the coding structure and utilizes its ability to understand global semantic information for feature extraction. The segmentation of small-scale target features is refined by the designed information grouping reconstruction convolution (IGRM) and channel classification reconstruction convolution (CRRM), which classify and reconstruct the extracted features by the amount of information. Finally, by integrating the up-sampling and down-sampling connections, the reconstructed features are fused with the features extracted by the encoder to form a multi-scale feature aggregation block to output the segmentation results. The refined reconstruction of small-scale target features is realized in multi-target scenarios with complex backgrounds, and high-quality segmentation maps are generated to improve the segmentation accuracy. Experimental results on the ISPRS Potsdam and ISPRS Vaihingen datasets show that the average intersection and merger ratio (mIoU) is 87.15% and 82.93%, respectively, and the overall accuracy (OA) is 91.53% and 91.4%, respectively. To verify the generalization ability of the model for small-scale target feature extraction in multi-target categories, this paper also designs a comparative experiment for the category of carts in complex backgrounds. The experimental results show that the mIoU on the UAVid dataset reaches 67.86%.
LÜJ, SHENQ, LÜM, et al. Research progress on semantic segmentation of remote sensing images based on deep learning[J]. Frontiers in Ecology and Evolution, 2023, 11: 1201125. (in Chinese)
LIUG Y, CAOY, ZENGZ Y,et al. Underwater multi-object segmentation technology based on spectral clustering with multi-feature weighting[J].Journal of Hunan University (Natural Sciences),2022,49(10):51-60.(in Chinese)
[5]
KUMARD, KUMARD .Hyperspectral image classification using deep learning models:a review[J]. Journal of Physics:Conference Series, 2021, 1950(1): 012087.
[6]
LONGJ, SHELHAMERE, DARRELLT .Fully convolutional networks for semantic segmentation[C]//2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Boston,MA,USA. IEEE,2015:3431-3440.
[7]
CHENL C, PAPANDREOUG, KOKKINOSI,et al .DeepLab:semantic image segmentation with deep convolutional nets,atrous convolution,and fully connected CRFs[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2018,40(4):834-848.
[8]
CHENL C, ZHUY K, PAPANDREOUG,et al. Encoder-decoder with atrous separable convolution for semantic image segmentation[M]//Computer Vision-ECCV 2018. Cham: Springer International Publishing,2018:833-851.
[9]
FUJ, LIUJ, TIANH J,et al .Dual attention network for scene segmentation[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Long Beach,CA,USA. IEEE, 2019: 3141-3149.
[10]
YUC Q, WANGJ B, PENGC,et al. BiSeNet:bilateral segmentation network for real-time semantic segmentation[M]//Computer Vision-ECCV 2018. Cham:Springer International Publishing, 2018: 334-349.
[11]
LIUZ, LINY T, CAOY, et al.Swin transformer:hierarchical vision transformer using shifted windows[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV). Montreal,QC, Canada. IEEE, 2021: 9992-10002.
[12]
WOO S, PARKJ, LEEJ Y, et al. CBAM:convolutional block attention module[M]//Computer Vision-ECCV 2018. Cham:Springer International Publishing, 2018: 3-19.
[13]
HUANGZ L, WANGX G, HUANGL C,et al .CCNet:criss-cross attention for semantic segmentation[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). Seoul,Korea (South). IEEE, 2019: 603-612.
[14]
CHENY P, FANH Q, XUB,et al .Drop an octave:reducing spatial redundancy in convolutional neural networks with octave convolution[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). Seoul,Korea (South). IEEE,2019:3435-3444.
[15]
CHOLLETF .Xception:deep learning with depthwise separable convolutions[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu,HI,USA. IEEE,2017:1800-1807.
WUY X, HEK M. Group normalization[C]// Computer Vision-ECCV 2018.Cham:Springer International Publishing,2018.
[18]
KRIZHEVSKYA, SUTSKEVERI, HINTONG E.ImageNet classification with deep convolutional neural networks[J].Communications of the ACM, 2017, 60(6): 84-90.
[19]
HUAB S, TRANM K, YEUNGS K .Pointwise convolutional neural networks[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City,UT,USA. IEEE,2018: 984-993.
[20]
LIR, ZHENGS Y, ZHANGC,et al .ABCNet:attentive bilateral contextual network for efficient semantic segmentation of fine-resolution remotely sensed imagery[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2021, 181: 84-98.
[21]
JIANGK X, LIUJ, ZHANGW H,et al. MANet:an efficient multidimensional attention-aggregated network for remote sensing image change detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 3328334.
[22]
WANGL B, LIR, ZHANGC,et al. UNetFormer:a UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery[J]. ISPRS Journal of Photogrammetry and Remote Sensing,2022, 190: 196-214.
[23]
LIR, ZHENGS Y, DUANC X,et al .Multistage attention ResU-net for semantic segmentation of fine-resolution remote sensing images[J].IEEE Geoscience and Remote Sensing Letters, 2021,19: 8009205.
[24]
STRUDELR, GARCIAR, LAPTEVI,et al. Segmenter:transformer for semantic segmentation[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV). Montreal, QC, Canada. IEEE,2021: 7242-7252.
[25]
DIAKOGIANNISF I, FURBYS, CACCETTAP,et al .SSG2:a new modelling paradigm for semantic segmentation[EB/OL]. 2023: 2310.08671.
[26]
ELHAJK, ALSHAMSID, ALDAHANA. GeoZ:a region-based visualization of clustering algorithms[J].Journal of Geovisua- lization and Spatial Analysis, 2023, 7(1): 15.
[27]
HONGX, ROOSEVELTC H. Orthorectification of large datasets of multi-scale archival aerial imagery:a case study from türkiye[J]. Journal of Geovisualization and Spatial Analysis,2023,7(2): 23.