This paper proposes a new key point detection algorithm, Point-GAT, addressing the challenge of suboptimal key point detection results at a small scale and the oversight of category semantic information between key points in previous methods. This algorithm joins fast connections to solve the learning degradation problem caused by a too deep network.It enhances the detection effect of small-scale objects by using deconvolution and feature fusion on Hourglass and ResNeXt backbone networks. Simultaneously, the algorithm employs the graph classification method to depict the semantic relationship between categories. This is achieved by constructing a directed weighted graph to capture the category semantic information among key points. During the optimizatipon of the positioning and regression function, the branch of the classification loss function is incorporated to reflect the category semantic information. The experimental results indicate that the average accuracy of this algorithm on the COCO dataset reaches 48.3%, and the average accuracy on the PASVAL VOC 2007 and PASVAL VOC 2012 datasets is higher than that of other algorithms.
目标检测分为两个子问题,即确定对象的中心和预测四个边界框。当预测对象的中心时,中心预测被集成到类别预测的目标中,并且还可以预测一个中心度得分。对于四个边界的预测是预测像素点与真实框的四个边之间的距离。具体的实现算法分为以下两种。 1)基于多关键点联合表达的方法:CornerNet-Lite使用左上角点和右下角点作为关键点进行目标检测;ExtremeNet[3]使用上下左右4个极值点和中心点作为关键点进行目标检测;CenterNet:Keypoint Triplets for Object Detection[4]使用左上角点、右下角点和中心点作为关键点;RepPoints[5]使用9个学习到的自适应跳动的采样点作为关键点;FoveaBox[6]使用中心点、左上角点和右下角点实现无anchor操作;PLN使用4个角点和中心点实现检测。2)基于单中心点检测的方法:CenterNet:Objects as Points使用中心点、宽度和高度作为检测的关键点;CSP[7]使用中心点和高度进行检测;FCOS[8]使用中心点和到框的2个距离实现关键点检测。
LIUL, OUYANGW L, WANGX G, et al. Deep learning for generic object detection: A survey [J]. International Journal of Computer Vision, 2020, 128(2): 261-318. DOI:10.1007/s11263-019-01247-4 .
[2]
JIAOL C, ZHANGF, LIUF, et al. A survey of deep learning-based object detection [J]. IEEE Access, 2019, 7(3): 128837-128868. DOI:10.1109/ACCESS.2019.2939201 .
[3]
LAWH, TENGY, RUSSAKOVSKYO, et al. CornerNet-lite: Efficient keypoint based object detection[EB/OL]. 2019: arXiv: 1904.08900.
[4]
ZHOUX, ZHUOJ C, KRÄHENBÜHLP. Bottom-up object detection by grouping extreme and center points [J]. Proceedings of the IEEE Computer Society Conference on Computer Vision, 2019,15(5): 850-859.
[5]
DUANK W, BAIS, XIEL X, et al. CenterNet:Keypoint triplets for object detection [J]. Proceedings of the IEEE International Conference on Computer Vision, 2019, 67(2): 6568-6577.
[6]
YANGZ, LIUS H, HUH, et al. RepPoints: Point set representation for object detection[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE Press,2020: 9656-9665. DOI:10.1109/ICCV.2019.00975 .
WANGC Y, LIAOH, WUY H, et al. CSPNet: A new backbone that can enhance learning capability of CNN[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops(CVPRW). New York: IEEE Press, 2020: 390-391. DOI:10.48550/arXiv.1911.11929 .
[9]
TIANZ, SHENC H, CHENH,et al. FCOS: Fully convolutional one-stage object detection [C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE Press, 2020: 657-666. DOI: 10.48550/arXiv.1904.01355 .
[10]
WUZ H, PANS R, CHENF W, et al. A comprehensive survey on graph neural networks [J]. IEEE Transactions on Neural Networks and Learning Systems, 2021, 32(1): 4-24. DOI:10.1109/TNNLS.2020.2978386 .
[11]
LIJ, RONGY, CHENGH, et al. Semi-supervised graph classification: A hierarchical graph perspective [C]//WWW'19: The World Wide Web Conference. New York: ACM, 2019: 972-982. DOI:10.1145/3308558.3313461 .
[12]
LIS D, GUOB H, YANGX B. Study on financial credit information based on graph neural network[J]. Computer Science, 2021,48(4): 85-90.
[13]
MORRISC, RITZERTM, FEY M, et al. Weisfeiler and leman go neural: Higher-order graph neural networks[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2019, 33: 4602-4609. DOI:10.1609/aaai.v33i01.33014602 .
[14]
GIRSHICKR, DONAHUEJ, DARRELLT, et al. Rich feature hierarchies for accurate object detection and semantic segmentation[C]//2014 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2014: 580-587. DOI:10.1109/CVPR.2014.81 .
[15]
XIES N, GIRSHICKR, DOLLÁRP, et al. Aggregated residual transformations for deep neural networks[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2017: 5987-5995. DOI:10.1109/CVPR.2017.634 .
[16]
SUNK, XIAOB, LIUD, et al. Deep high-resolution representation learning for human pose estimation[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press,2019 : 5686-5696. DOI:10.1109/CVPR.2019.00584 .
[17]
CHENGB W, XIAOB, WANGJ D, et al. HigherHRNet: Scale-aware representation learning for bottom-up human pose estimation [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press,2020 : 5385-5394. DOI:10.1109/CVPR42600.2020.00543 .
[18]
HEK M, GKIOXARIG, DOLLÁRP, et al. Mask R-CNN[C]//2017 IEEE International Conference on Computer Vision. New York: IEEE Press, 2017: 2980-2988. DOI:10.1109/ICCV.2017.322 .
[19]
LIUS, QIL, QINH F, et al. Path aggregation network for instance segmentation[C]// 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 8759-8768. DOI:10.1109/CVPR.2018.00913 .
[20]
LIY H, CHENY T, WANGN Y, et al. Scale-aware trident networks for object detection[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE Press, 2019: 6053-6062. DOI:10.1109/ICCV.2019.00615 .
[21]
REDMONJ, FARHADIA. YOLO9000: Better, faster, stronger[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2017: 6517-6525. DOI:10.1109/CVPR.2017.690 .
[22]
LINT Y, GOYALP, GIRSHICKR, et al. Focal loss for dense object detection [C]//2017 IEEE International Conference on Computer Vision. New York: IEEE Press,2017 : 2999-3007. DOI:10.1109/ICCV.2017.324 .
[23]
LAWH, DENGJ. CornerNet: Detecting objects as paired keypoints[C]//Computer Vision―ECCV 2018. Cham: Springer International Publishing, 2018: 765-781. DOI:10.1007/978-3-030-01264-9_45 .
[24]
ZHUC C, HEY H, SAVVIDESM. Feature selective anchor-free module for single-shot object detection[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2020: 840-849. DOI:10.1109/CVPR.2019.00093 .
[25]
KEW, CHENJ, JIAOJ B, et al. SRN: Side-output residual network for object symmetry detection in the wild[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press,2017 : 302-310. DOI:10.1109/CVPR.2017.40 .
[26]
HEK M, ZHANGX Y, RENS Q, et al. Identity mappings in deep residual networks[C]//Computer Vision―ECCV 2016. Cham: Springer International Publishing, 2016: 630-645. DOI:10.1007/978-3-319-46493-0_38 .
[27]
GEW F, YANGS B, YUY Z. Multi-evidence filtering and fusion for multi-label classification, object detection and semantic segmentation based on weakly supervised learning[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 1277-1286. DOI:10.1109/CVPR.2018.00139 .
[28]
FUJ L, ZHENGH L, MEIT. Look closer to see better: Recurrent attention convolutional neural network for fine-grained image recognition[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2017: 4476-4484. DOI:10.1109/CVPR.2017.476 .
[29]
LIUW, ANGUELOVD, ERHAND, et al. SSD: Single shot MultiBox detector[C]//Computer Vision-2016. Cham: Springer International Publishing, 2016: 21-37. DOI:10.1007/978-3-319-46448-0_2 .
[30]
REDMONJ, DIVVALAS, GIRSHICKR, et al. You only look once: Unified, real-time object detection[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press,2016: 779-788. DOI:10.1109/CVPR.2016.91 .