Key Laboratory of Aerospace Information Security and Trusted Computing,Ministry of Education,School of Cyber Science and Engineering,Wuhan University,Wuhan 430072,Hubei,China
The task of image geo-localization refers to the prediction of geographic label for given image. Such prediction was achieved by matching the input with existing geographic labelled image in the database. Due to the lack of ground images with geographic labels, databases were often established by satellite images. However, the huge perspective difference between satellite images and ground images brought great challenges to image matching. In this paper, we proposed a new conditional generative adversarial network (CGAN) called Crossview Attention Seq (CAS) for cross-view image conversion. CAS achieved better effect through semantic segmentation of image, meanwhile we introduced spatial attention mechanism to suppress the noise in transformation. The auxiliary information was input to the framework to optimize the parameters. The image matching framework was built based on Siamese network architecture. In addition, the improved loss function was integrated into the training process. Compared with Triplet loss, it greatly improved the optimization effect. Experimental results prove the effectiveness and superiority of our method on two classical datasets. Compared with the existing methods, our model is more accurate in geo-location and more robust to low quality data.
ZITOVÁB, FLUSSERJ. Image registration methods: A survey [J]. Image and Vision Computing, 2003, 21(11): 977-1000. DOI:10.1016/S0262-8856(03)00137-9 .
[2]
VON N, HAYSJ. Localizing and Orienting Street Views Using Overhead Imagery [EB/OL]. [2022-01-02]. DOI: 10.1007/978-3-319-46448-0_30 .
[3]
HUS X, FENGM D, NGUYENR M H, et al. CVM-net: Cross-view matching network for image-based ground-to-aerial geo-localization [C]// Proceedings of the Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 7258-7267. DOI:10.1109/CVPR.2018.00758 .
[4]
BANSALM, SAWHNEYH S, CHENGH, et al. Geo-localization of street views with aerial image databases [C]// Proceedings of the 19th ACM International Conference on Multimedia. New York: ACM, 2011: 1125-1128. DOI:10.1145/2072298.2071954 .
[5]
VISWANATHANA, PIRESB R, HUBERD. Vision based robot localization by ground to satellite matching in GPS-denied situations [C]//2014 IEEE/RSJ International Conference on Intelligent Robots and Systems. New York: IEEE Press, 2014: 192-198. DOI:10.1109/IROS.2014.6942560 .
[6]
JÉGOUH, DOUZEM, SCHMIDC, et al. Aggregating local descriptors into a compact image representation [C]//2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2010: 3304-3311. DOI:10.1109/CVPR.2010.5540039 .
[7]
ARANDJELOVICR, GRONATP, TORIIA, et al. NetVLAD: CNN architecture for weakly supervised place recognition [C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2016: 5297-5307. DOI: 10.1109/TPAMI.2017.2711011 .
[8]
CAIS, GUOY, KHANS, et al. Ground-to-aerial image geo-localization with a hard exemplar reweighting triplet loss [C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. New York: IEEE Press, 2019: 8391-8400. DOI: 10.1109/ICCV.2019.00848 .
[9]
GOODFELLOWI, POUGET-ABADIEJ, MIRZAM, et al. Generative adversarial nets [J]. Advances in Neural Information Processing Systems, 2014, 27(2): 2672-2680. DOI: 10.5555/2969033.2969125 .
HERMANSA, BEYERL, LEIBEB. In Defense of the Triplet Loss for Person Re-identification[EB/OL]. [2021-10-24]. DOI: 10.48550/arXiv.1703.07737 .
[12]
CHENW H, CHENX T, ZHANGJ G, et al. Beyond triplet loss: A deep quadruplet network for person re-identification [C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2017: 403-412. DOI: 10.1109/CVPR.2017.145 .
[13]
LIUL, LIH D, DAIY C. Stochastic attraction-repulsion embedding for large scale image localization [C]// Proceedings of the International Conference on Computer Vision. New York: IEEE Press, 2019: 2570-2579. DOI:10.1109/ICCV.2019.00266 .
[14]
ZHAIM H, BESSINGERZ, WORKMANS, et al. Predicting ground-level scene layout from aerial imagery[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2017: 867-875. DOI: 10.1109/CVPR.2017.440 .
[15]
REGMIK, BORJIA. Cross-view image synthesis using conditional GANs [C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 3501-3510. DOI: 10.1109/CVPR.2018.00369 .
[16]
WORKMANS, SOUVENIRR, JACOBSN. Wide-area image geolocalization with aerial reference imagery [C]//Proceedings of the IEEE International Conference on Computer Vision. New York: IEEE Press, 2015: 3961-3969. DOI: 10.1109/ICCV.2015.451 .
[17]
LING S, MILANA, SHENC H, et al. Refinenet: Multi-path refinement networks for high-resolution semantic segmentation [C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2017: 5168-5177. DOI: 10.1109/CVPR.2017.549 .
[18]
SIMONYANK, ZISSERMANA. Very Deep Convolutional Networks for Large-scale Image Recognition [EB/OL]. [2021-10-23]. DOI: 10.48550/arXiv.1409.1556 .
[19]
ROFFOG, MELZIS, CASTELLANIU, et al. Infinite feature selection: A graph-based feature filtering approach [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, 43(12): 4396-4410. DOI: 10.1109/TPAMI.2020.3002843 .
[20]
XIAS Y, ZHANGH, LIW H, et al. GBNRS: A novel rough set algorithm for fast adaptive attribute reduction in classification [J]. IEEE Transactions on Knowledge and Data Engineering, 2022, 34(3):1231-1242. DOI: 10.1109/TKDE.2020.2997039 .
[21]
SUNB, CHENC, ZHUY Y, et al. Geocapsnet: Aerial to Ground View Image Geo-Localization Using Capsule Network [EB/OL]. [2021-11-02]. DOI: 10.48550/arXiv.1904.06281 .