The existing human pose estimation methods ignore the calculation load of terminal device to achieve high detection accuracy when dealing with complex scenes and multi-objective tasks. To address the issues, this paper proposes a lightweight dual-path multi-scale feature fusion pose estimation network model. The model adds a high-resolution branch to the Litepose network and combines multiple dilation-based bottleneck convolutions (Multi-dilated Block, MDBlock) to form a bottleneck block for feature extraction, while preserving as much detail information as possible. Secondly, by constructing a lightweight and efficient feature extraction unit, a multi-branch efficient multi-scale attention mechanism (MEMA) is proposed, which enhances the network's attention to different regions. After the final up-sampling operation, a dual pooling attention feature fusion mechanism is designed. It can simultaneously pay attention to global and local channel feature information and achieve strong feature fusion. Experimental results show that the proposed method, compared with the Litepose method, only incurs a small increase in computational cost and parameter count on the COCO2017 database and a more challenging database CrowdPose. The average mAPs on these two databases are improved by 7.6% and 4.4%, respectively.
多重扩张瓶颈卷积分支的输出在通道特征的矩阵乘法操作之前转换为相应的维度形状。通过并行处理得到的输出做矩阵乘法,得出第一个空间注意力图,它在同一阶段收集不同尺度的空间信息。此外,同样利用2D全局平均池化对多重扩张瓶颈卷积分支中的全局空间信息进行编码,并且1×1分支在通道特征的矩阵乘法操作之前转换为相应的维度形状,之后,得到第二个空间注意力图,它保留了准确的全局空间位置信息。最后,各分支的输出特征图被计算为两个生成的空间注意力权重值的聚合,经过Sigmoid函数赋值给最初输入的特征张量。经过MEMA得到的最终输出与输入特征图 X 尺寸大小相同,通过两个分支之间的跨空间学习,捕获像素级成对关系并突出显示所有像素的全局上下文信息,促使模型学习到更丰富、更有区分性的特征表示。
LiuW, BaoQ, SunY, et al. Recent advances of monocular 2D and 3D human pose estimation: a deep learning perspective[J]. ACM Computing Surveys, 2022, 55(4): 550480.
SunZhi-yong, LiHong-you, YeJun-yong. 3D human joint point recognition based on weakly supervised migration networks[J]. Journal of Jilin University (Engineering and Technology Edition), 2024, 54(1): 251-258.
[4]
GengZ, SunK, XiaoB, et al. Bottom-up human pose estimation via disentangled keypoint regression[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2021: 14676-14686.
[5]
QiY F, ZhangH R, LiuJ. More accurate heatmap generation method for human pose estimation[J]. Multimedia Systems, 2024, 30(4): 180.
WangYu, ZhaoKai. Post-processing of human posture thermograms based on sub-pixel localization[J]. Journal of Jilin University (Engineering and Technology Edition), 2024, 54(5): 1385-1392.
[8]
CaoZ, SimonT, WeiS, et al. Realtime multi-person 2d pose estimation using part affinity fields[C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, USA, 2017: 7291-7299.
[9]
SunK, XiaoB, LiuD, et al. Deep high-resolution representation learning for human pose estimation[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 5693-5703.
HaoHe-fei, ZhangLong-hao, CuiHong-zhen, et al. Deep neural networks in human pose estimation: a systemic review[J]. Computer Engineering and Applications, 2025,61(9): 41-60.
[12]
KimW, SungJ, SaakesD, et al. Ergonomic postural assessment using a new open-source human pose estimation technology(openpose)[J]. International Journal of Industrial Ergonomics, 2021, 84: 103164.
[13]
XuM, GuoL, WuH. Robust abnormal human-posture recognition using openpose and multiview cross-information[J]. IEEE Sensors Journal, 2023, 23(11): 12370-12379.
MengCai-xia, XueHong-qiu, ShiLei, et al. Openpose human fall detection algorithm based on attention mechanism[J]. Journal of Computer-Aided Design & Computer Graphics, 2024, 36(12): 2040-2050.
[16]
WangJ, SunK, ChengT, et al. Deep high-resolution representation learning for visual recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 43(10): 3349-3364.
[17]
YuC, XiaoB, GaoC, et al. Lite-hrnet: a lightweight high-resolution network[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2021: 10440-10450.
[18]
MaN, ZhangX, ZhengH T, et al. Shufflenet v2: practical guidelines for efficient cnn architecture design[C]∥Proceedings of the European Conference on Computer Vision, Munichi, Germany, 2018: 116-131.
[19]
WangY, LiM, CaiH, et al. Lite pose: efficient architecture design for 2d human pose estimation[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, USA, 2022: 13126-13136.
[20]
HowardA G. Mobilenets: efficient convolutional neural networks for mobile vision applications[J]. Arxiv Preprint, 2017,arXiv:
[21]
HouQ, ZhouD, FengJ. Coordinate attention for efficient mobile network design[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2021: 13713-13722.
[22]
SamadhJ H A, KhatibS K A. To perceive or not to perceive: lightweight stacked hourglass network[J]. Arxiv Preprint, 2023,arXiv:
[23]
DaiY, GiesekeF, OehmckeS, et al. Attentional feature fusion[C]∥Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, USA, 2021: 3560-3569.
[24]
LinT, MaireM, BelongieS, et al. Microsoft coco: common objects in context[C]∥13th European Conference on Computer Vision, Zurich, Switzerland, 2014: 740-755.
[25]
LiJ, WangC, ZhuH, et al. Crowdpose: efficient crowded scenes pose estimation and a new benchmark[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 10863-10872.
[26]
ChengB, XiaoB, WangJ, et al. Higherhrnet: scale-aware representation learning for bottom-up human pose estimation[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 5386-5395.
[27]
XuY, ZhangJ, ZhangQ, et al. Vitpose: simple vision transformer baselines for human pose estimation[J]. Advances in Neural Information Processing Systems, 2022, 35: 38571-38584.
[28]
GengZ, SunK, XiaoB, et al. Bottom-up human pose estimation via disentangled keypoint regression[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2021: 14676-14686.
[29]
MajiD, NagoriS, MathewM, et al. Yolo-pose: enhancing yolo for multi person pose estimation using object keypoint similarity loss[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, USA, 2022: 2637-2646.
[30]
JiangT, LuP, ZhangL, et al. Rtmpose: real-time multi-person pose estimation based on mmpose[J]. Arxiv Preprint, 2023, arXiv:
[31]
WangH, LiuJ, TangJ, et al. Lightweight super-resolution head for human pose estimation[C]∥Proceedings of the 31st ACM International Conference on Multimedia, Ottawa, Canada, 2023: 2353-2361.
[32]
LuP, JiangT, LiY, et al. Rtmo: towards high-performance one-stage real-time multi-person pose estimation[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2024: 1491-1500.