To address the issue that human motion image acquisition is susceptible to environmental variations (such as illumination, occlusion, and background complexity), which can lead to difficult pose feature extraction and reduced pose capture accuracy, a human motion pose information capture method based on multi-feature fusion is proposed. A dual-branch architecture is designed for the human motion pose information capture model. In the upper branch, the current frame together with its two adjacent preceding and subsequent frames is taken as the input to a Transformer-based video frame temporal feature extraction network, by which inter-frame temporal features are learned and pose variation information is captured. In the lower branch, skeleton data of video frames is obtained using OpenPose, and with the center of the pelvic region as the origin, the data is transformed into a spherical coordinate system to construct a spherical descriptor of joints (SDJ) matrix. This matrix is then used as the input to a CNN-based motion pose feature network, where motion pose features are extracted. These features, together with the temporal features, are subsequently fed into a complementary feature fusion network to achieve complementary integration of human motion pose features, and human motion pose probabilities are predicted through a Softmax layer, thereby enabling human motion pose information capture. Experimental results show that when the video frame sequence length is set to 5, an F1 score of approximately 0.94 is achieved. The proposed method effectively captures human motion pose information, consistently achieving higher PCK values than the compared methods, with an MPJPE of 28.5 and an AUC of 0.923.
MaYa-tong, WangSong, LiuYing-fang. Research on Human Action Recognition Method by Fusing Multimodal Data[J]. Computer Engineering, 2022, 48(9):180-188.
MaXiao, YanYu-dong. Multiscale spatio-temporal correlation feature learning for human pose estimation [J].Journal of South-Central University for Nationalities (Natural Science Edition), 2023, 42(1): 95-102.
FengXin-xin, LiWen-long, HeZhao,et al.Human posture recognition based on multi-dimensional information feature fusion of frequency modulated continuous wave radar[J]. Journal of Electronics & Information Technology, 2022, 44(10): 3583-3591.
SuBen-yue, ZhangPeng, ZhuBang-guo, et al. Human action recognition based on skeleton edge information under projection subspace[J]. Journal of System Simulation, 2024, 36(3): 555-563.
LeiYong-sheng, DingMeng, ShenYao, et al. Action recognition model based on improved two stream vision transformer[J]. Computer Science, 2024, 51(7): 229-235.
CaoJian-rong, LvJun-jie, WuXin-ying, et al. Fall detection algorithm integrating motion features and deep learning[J]. Journal of Computer Applications, 2021, 41(2): 583-589.
WuZi-yi, ChenMin-rong. Multi-stream convolutional human action recognition based on the fusion of spatio-temporal domain attention module[J]. Journal of South China Normal University (Natural Science Edition), 2023, 55(3): 119-128.
TianZhi-qiang, DengChun-hua, ZhangJun-wen. Human behavior recognition algorithm based on skeletal temporal divergence feature[J]. Journal of Computer Applications, 2021, 41(5): 1450-1457.
JiChen-zhong, Ci Wang jin-mei, ZhangWei, et al. Research on action recognition based on improving 2D CNN spatial-temporal feature extraction[J]. Journal of Chinese Computer Systems, 2024,45(1):168-176.
ZhangCong-cong, HeNing, SunQi-xiang, et al. Human motion recognition method based on attention mechanism of 3D DenseNet [J]. Computer Engineering, 2021, 47(11): 313-320.