Aiming at the problem that the target occlusion leads to the loss of some target information of the image in human pose estimation, a light-weight and high-resolution cascaded pyramid network model was constructed based on HRNet-32 to estimate the occlusion body pose. The improved Gaff module of GhostNet is introduced into the first stage of high-resolution network (HRNet-32). The network is light-weight, features are extracted initially and multi-scale feature fusion training is carried out. A cascaded pyramid network (CPN) was added to HRNet-32 for secondary feature extraction to obtain the key points of the occluded part of the human body, and a regression heat map was used to estimate the body posture. Experimental results on public data sets MPII and 3DOH50K show that the accuracy of the proposed model is better than that of HRNet-32.
表1是本文方法与其他方法的参数量和运算复杂度的结果比较。HRNet-32作为基础框架网络,在模型参数设置上采用了线性变换缩放系数s=2来进行冗余特征图的生成,根据(1)式和(2)式计算,HRNet-32 with Gaff module相比CPN方法[6],运算复杂度降低1.1%,参数量降低31.5%;相比HRNet-32 without Gaff module,运算复杂度降低13.6%,参数量降低35.1%;相比SimpleBaseline[6]方法,运算复杂度降低31.1%,参数量降低45.6%。本文的网络模型相较于HRNet-32 without Gaff module+CPN,网络运算复杂度和参数量分别降低了2 GB和9.8 MB。
表2展示了在MPII数据集上,将以头部作为归一化参数的关键点正确估计的比例PCKh (percentage of correct keypoints based on head)作为评估标准,评估的关键点为头部(Head)、肩部(Sho.)、肘部(Elb.)、腕部(Wri.)、髋部(Hip)、膝部(Knee)、踝部(Ank.),在不同的网络上训练的实验结果。相比于文献[14]提出的CPMs、文献[15]的对抗性PoseNet、文献[16]的PRMs、文献[17]的多尺度结构感知神经网络、文献[18]的DLCM(deeply learned compositional models)和文献[7]的DHRRL(deep high-resolution representation learning)方法,本文提出的GHRCPN网络在MPII数据集上对大部分关键点能获得最好的识别结果,并且平均精度达到了93.4%。这是由于对HRNet-32改进后融入CPN提取了被遮挡部分的关键特征。
表3为在MPII数据集上进行消融实验的结果。将改进得到的Gaff模块融入GhostNet,其平均估计精确度相对于GhostNet without AFF module有0.8%的提升。将改进得到的Gaff模块融入到高分辨率网络HRNet-32,平均估计精确度比无Gaff 模块的HRNet-32平均估计精确度提高了0.6%。而在HRNet-32 with Gaff module 上融入CPN后所得到的网络的平均估计精确度比HRNet-32 without Gaff module和HRNet-32 with Gaff module分别有1.1%和0.5%的提升。
TAOY, YANGF, LIUY,et al. Research and analysis of K-means clustering algorithm [C]// Proceedings of the 2016 Academic Annual Meeting of Guangxi Computer Society. Guangxi: Guangxi Computer Society,2016: 2-6 (Ch) . DOI: 10.21436/inbom.12382432 .
[5]
TOSHEVA, SZEGEDYC. DeepPose: Human pose estimation via deep neural networks [C]//2014 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2014: 1653-1660. DOI:10.1109/CVPR.2014.214 .
[6]
LINT Y, DOLLÁRP, GIRSHICKR, et al. Feature pyramid networks for object detection[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2017: 936-944. DOI:10.1109/CVPR.2017.106 .
[7]
NEWELLA, YANGK Y, DENGJ. Stacked hourglass Networks for human pose estimation[J]. European Conference on Computer Vision.Cham:Springer, 2016:483-499. DOI:10.1007/978-3-319-46484-8_29 .
[8]
CHENY L, WANGZ C, PENGY X, et al. Cascaded pyramid network for multi-person pose estimation[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). New York: IEEE Press, 2018: 7103-7112. DOI:10.1109/CVPR.2018.00742 .
[9]
SUNK, XIAOB, LIUD, et al. Deep high-resolution representation learning for human pose estimation[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2019: 5686-5696. DOI:10.1109/CVPR.2019.00584 .
[10]
HANK, WANGY H, TIANQ, et al. GhostNet: more features from cheap operations[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2020: 1577-1586. DOI:10.1109/CVPR42600.2020.00165 .
[11]
DAIY, GIESEKEF, OEHMCKES, et al. Attentional Feature Fusion[EB/OL].[2020-09-10].DOI: 10.1109/wacv48630.2021.00360 .
[12]
CHUX, OUYANGW L, LIH S,et al. Structured feature learning for pose estimation[C]// 2016 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR),2016:4715-4723. DOI: 10.1109/CVPR.2016.510 .
LIUP K, ZHUC J, ZHANGY. Research on improved lightweight high resolution human keypoint detection [J]. Computer Engineering and Applications, 2021,57(2): 143-149. DOI:10.3778/j.issn.1002-83 31. 2007-0276(Ch ).
[15]
ZHANGT S, HUANGB Z, WANGY G. Object-occluded human shape and pose estimation from a single color image[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) New York: IEEE Press, 2020: 7374-7383. DOI:10.1109/CVPR42600.2020.00740 .
[16]
ANDRILUKAM, PISHCHULINL, GEHLERP, et al. 2D human pose estimation: New benchmark and state of the art analysis[C]//2014 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2014: 3686-3693. DOI:10.1109/CVPR.2014.471 .
[17]
WEIS H, RAMAKRISHNAV, KANADET, et al. Convolutional pose machines [C]//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2016: 4724-4732. DOI:10.1109/CVPR.2016.511 .
[18]
CHENY, SHENC H, WEIX S, et al. Adversarial PoseNet: A structure-aware convolutional network for human pose estimation[C]//2017 IEEE International Conference on Computer Vision (ICCV). New York: IEEE Press, 2017: 1221-1230. DOI:10.1109/ICCV.2017.137 .
[19]
YANGW, LIS, OUYANGW L, et al. Learning feature Pyramids for human pose estimation[C]//2017 IEEE International Conference on Computer Vision (ICCV). New York: IEEE Press, 2017: 1290-1299. DOI:10.1109/ICCV.2017.144 .
[20]
KEL, CHANGM C, QIH, et al. Multi-scale structure-aware network for human pose estimation [C]//European Conference on Computer Vision. Cham: Springer,2018:731-746. DOI:10.1007/978-3-030-01216-8_44 .
[21]
TANGW, YUP, WUY. Deeply learned compositional models for human pose estimation [C]// European Conference on Computer Vision. Cham: Springer,2018:197-214. DOI:10.1007/978-3-030-01219-9_12 .
[22]
VINYALSO, TOSHEVA, BENGIOS, et al. Show and tell: Lessons learned from the 2015 MSCOCO image captioning challenge[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(4): 652-663. DOI:10.1109/TPAMI.2016.2587640 .
[23]
LIJ, SUW, WANGZ F. Simple pose: Rethinking and improving a bottom-up approach for multi-person pose estimation [J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2020, 34(7): 11354-11361. DOI:10.1609/aaai.v34i07.6797 .
[24]
NIEY L, LEEJ H, YOONS, et al. A multi-stage convolution machine with scaling and dilation for human pose estimation [J]. KSII Transactions on Internet and Information Systems, 2019,13(6):3182-3198. DOI:10.3837/tiis.2019.06.023 .
[25]
LIJ F, WANGC, ZHUH, et al. CrowdPose: Efficient crowded scenes pose estimation and a new benchmark[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2019: 10855-10864. DOI:10.1109/CVPR.2019.01112 .