Due to the large difference in the scale of the object, the arbitrary direction and the dense distribution of the object in the remote sensing image, the existing detection methods rarely pay direct attention to the dense edge information and the object cannot obtain a suitable receptive field, so it is difficult to have good detection results in remote sensing detection. In order to solve the above problems, this paper proposed a multi-scale feature enhancement network based on large kernel convolution and dense object refinement (LKCSFP-NET) for remote sensing image detection. Firstly, the network based on SKNET added a cavity convolution to form a large kernel convolution block (LKB) to obtain the best sensitivity field for small targets and improve the adaptability and accuracy of the network to multiple scales. Secondly, on the basis of FPN, the centralized spatial feature pyramid CSFP module was added to solve the problem of low detection efficiency of remote sensing images due to dense object distribution and complex detection background by combining global semantic information with local semantic information. The experimental results show that on the DOTA and HRSC2016 public datasets, the average detection accuracy of the proposed algorithm on the two datasets is 74.90% and 96.60%, respectively, which is 1.36 and 0.63 percentage points higher than that of the baseline network, which is better than most existing models. The proposed LKCSFP-NET has stable performance in the two public datasets, and has good detection results for small objects and densely arranged objects, which is higher than the detection accuracy of most existing models, and can be well applied to the detection of remote sensing objects.
WANGQ, GUOJ Y, YUANY. Embedding Structured Contour and location prior in siamesed fully convolutional networks for road detection[J].IEEE Transactions on Intelligent Transportation Systems, 2018, 19(1): 230-241 .
[2]
XIEW Y, LEIJ, FANGS, et al. Dual feature extraction network for hyperspectral image analysis[J]. Pattern Recognition,2021, 118: 107992.
[3]
XIEW Y, LEIJ, CUIY H, et al. Hyperspectral pansharpening with deep priors[J]. IEEE Transactions on Neural Networks and Learning Systems, 2020, 31(5): 1529-1543.
[4]
GANCIG, CAPPELLOA, BILOTTAG, et al. How the variety of satellite remote sensing data over volcanoes can assist hazard monitoring efforts: The 2011 eruption of Nabro Volcano[J]. Remote Sensing of Environment, 2020, 236: 111426.
[5]
XIEW Y, ZHANGX, LIY S, et al. Weakly supervised low-rank representation for hyperspectral anomaly detection[J]. IEEE Transactio- ns on Cybernetics, 2021, 51(8): 3889–3900.
[6]
WANGQ, GAOJ Y, YUANY. A joint convolutional neural networks and context transfer for street scenes labeling[J]. IEEE Transactions on Intelligent Transportation Systems, 2018, 19(5): 1457-1470.
[7]
CEHNGG, HANJ W, ZHOUP C, et al. Multi-class geospatial object detection and geographic image classification based on collection of part detectors[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2014, 98: 119-132.
[8]
LIK, WANG, CHENGG, et al. Object detection in optical remote sensing images: A survey and a new benchmark[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2020, 159: 296-307.
[9]
XIAG S, BAIX, DINGJ, et al. Dota: A large-scale dataset for object detection in aerial images[C]//IEEE Conference on Computer Vision and Pattern Recognition, 2018: 3974-3983.
[10]
LIUZ K, WANGH Z, WENGL B, et al. Ship rotated bounding box space for ship extraction from high-resolution optical satellite images with complex backgrounds[J]. IEEE Geoscience and Remote Sensing Letters, 2016, 13(8): 1074-1078.
[11]
DINGJ, XUEN, LONGY, et al. Learning roi transformer for oriented object detection in aerial images[C]//IEEE Conference on Computer Vision and Pattern Recognition, 2019: 2844-2853.
[12]
LIUZ K, HUJ G, WENGL B, et al. Rotated region based cnn for ship detection[C]//IEEE International Conference on Image Procecssing, 2017: 900-904.
[13]
QUANY, ZHANGD, ZHANGL Y, et al. Centralized feature pyramid for object detection[J]. IEEE Transactions on Image Processing, 2023, 32: 4341-4354.
[14]
LINT Y, DOLLÁRP, GIRSHICKR, et al. Feature pyramid networks for object detection[C]//IEEE Conference on Computer Vision and Pattern Recognition, 2017: 936-944.
[15]
LIUS, QIL, QINH F, et al. Path aggregation network for instance segmentation[C]//IEEE Conference on Computer Vision and Pattern Recognition, 2018: 8759-8768.
[16]
GUOC X, FANB, ZHANGQ, et al. AugFPN: Improving multi-scale feature learning for object detection[C]//IEEE Conference on Computer Vision and Pattern Recognition, 2020: 12592-12601.
[17]
LIS, YANGL X, HUANGJ Q, et al. Dynamic anchor feature selection for single-shot object detection[C]//IEEE International Conference on Computer Vision, 2019: 6608-6617.
[18]
LIX, WANGW H, HU X LM, et al. Selective kernel networks[C]//IEEE Conference on Computer Vision and Pattern Recognition, 2019: 510-519.
[19]
TENGZ, DUANY N, LIUY, et al. Global to local: Clip-LSTM-based object detection from remote sensing images[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 5603113.
[20]
SHIL K, KUANGL Y, XUX, et al. CANet: Centerness-aware network for object detection in remote sensing images[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 5603613.
[21]
HANW, FANR Y, WANGLZ, et al. Improv- ing training instance quality in aerial image object detection with a sampling-balance-based multistage network[J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 59(12): 10575-10589.
[22]
WANGG Q, ZHUANGY, CHENGH, et al. FSoD-Net: Full-scale object detection from optical remote sensing imagery[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 5602918.
[23]
HUANGW, LIG Y, CHENQ Q, et al. CF2PN: A cross-scale feature fusion pyramid network based remote sensing target detection[J]. Remote Sensing, 2021, 13(5): 847.
[24]
XUT, SUNX, DIAOW H, et al. ASSD: Feature aligned single-shot detection for multiscale objects in aerial imagery[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 5607117.
[25]
YANGX, YANJ C, FENGZ M, et al. R3Det: Refined single-stage detector with feature refinement for rotating object[C]//AAAI Conference on Artificial Intelligence, 2021, 35(4): 3163-3171.
[26]
YANGX, YANJ C. Arbitrary-oriented object detection with circu-lar smooth label[C]//European Conference on Computer Vision, 2020: 677-694.
[27]
YANGX, HOUL P, ZHOUY, et al. Dense label encoding for boundary discontinuity free rotation detection[C]//IEEE Conference on Computer Vision and Pattern Recognition, 2021: 15814-15824.
[28]
YANGX, YANJ C, LIAOW L, et al. Scrdet++: Detecting small, cluttered and rotated objects via instance-level feature denoising and rotation loss smoothing[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(2): 2384-2399.
[29]
HANJ M, DINGJ, LIJ, et al. Align deep features for oriented object detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 5602511.
[30]
PANX, RENY Q, SHENGK K, et al. Dynamic refinement network for oriented and densely packed object detection[C]//IEEE Conference on Computer Vision and Pattern Recognition, 2020: 11204 -11213.
[31]
QIANW, YANGX, PENGS L, et al. Learning modulated loss for rotated object detection[C]//AAAI Conference on Artificial Intelligence, 2021: 2458-2466.
[32]
CHENGG, WANGJ B, LIK, et al. Anchor-free oriented proposal generator for object detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 5625411.
[33]
RanftlRené, BochkovskiyAlexey, KoltunVladlen. Vision transformers for dense prediction[C]//IEEE International Conference on Computer Vision, 2021: 12159-12168.
[34]
YANH T, LIZ, LIW J, et al. Contnet: Why not use convolution and transformer at the same time?[DB/OL]. (2021-05-10)[2024-01-03].
[35]
ZHENGS X, LUJ C, ZHAOH S, et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers[C]//IEEE Conference on Computer Vision and Pattern Recognition, 2021: 6877-6886.
[36]
LUOW J, LIY J, UrtasunRaquel, et al. Understanding the effective receptive field in deep convolutional neural networks[C]//International Conference on Neural Information Processing Systems, 2016: 4905-4913.
[37]
LIUZ, MAOH Z, WUC Y, et al. A convnet for the 2020s[C]//IEEE Conference on Computer Vision and Pattern Recognition, 2022: 11966-11976.
[38]
DINGX H, ZHANGX Y, HANJ G, et al. Scaling up your kernels to 31×31: Revisiting large kernel design in cnns[C]//IEEE Conference on Computer Vision and Pattern Recognition, 2022: 11953-11965.
[39]
LIUS W, CHENT L, CHENX H, et al. More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity[DB/OL]. (2023-03-03)[2024-01-03].