In order to solve the problem that the current fitness exercise evaluation systems provided poor evaluation information and involved high computational complexity, a fitness exercise evaluation system based on human posture estimation was proposed. The system utilized the built-in camera of mobile phones or tablet computers, combined with human pose estimation algorithms, action recognition algorithms, and action evaluation algorithms to achieve intelligent and effective evaluation of users’ fitness exercises. Firstly, the MediaPipe algorithm was used to estimate the human posture and obtain the human joint bone data.Then, it was fed into the proposed SCBT-GCN network for action recognition. Based on kinematic principles, a series of evaluation methods had been designed which evaluate joint angles, symmetry, trajectory characteristics, and movement fluency. Finally, corresponding targeted movement assessment algorithms were invoked based on the identified action categories to detect and evaluate abnormal fitness movements, thereby forming a fitness action evaluation system with strong intelligence and high real-time performance. Experimental results show that the system has an accuracy of 98.6% in action recognition, which can detect abnormal actions in real time and score fitness actions offline, demonstrating high practical value.
动作评估方法主要分为模板匹配和空间-时间法。模板匹配是指将待测动作序列与预先定义动作序列的模板进行比较。Song等[12]基于人体的关键点数据,设计距离函数,采用动态时间规整(dynamic time warping,DTW)算法来度量时间序列之间的相似性,实现动作评估。于景华等[13]设置了标准动作和实验动作,计算两个动作帧的平均帧距离,通过DTW将标准动作和实验动作之间的距离进行比较,归一化后获得动作相似度评分。韩丽等[14]通过计算向量的夹角余弦值,利用特征平面相似性匹配的方法计算运动数据相关匹配度,确定舞蹈表演者的姿势与标准姿势的相似程度。空间-时间法是指分析一个动作中不同身体部位的空间和时间的关系。Alexiadis等[15]基于Kinect深度相机采集的骨骼点坐标、关节角度和旋转量,提出了一组基于四元数相关的测量分数,与教练的标准动作进行对比,得到舞蹈动作评分。以上运动评估的研究中多使用普遍单一的评估方法,没有针对不同运动的发力特点设计不同方法,因此,评估方法的适应性还需进一步提高。
ToshevA, SzegedyC.DeepPose:human pose estimation via deep neural networks[C]//2014 IEEE Conference on Computer Vision and Pattern Recognition.Columbus:IEEE,2014:1653-1660.
[2]
NewellA, YangK Y, DengJ.Stacked hourglass networks for human pose estimation[C]//Computer Vision-ECCV 2016.Cham:Springer International Publishing,2016:483-499.
[3]
FangH S, XieS Q, TaiY W,et al.RMPE:regional multi-person pose estimation[C]//2017 IEEE International Conference on Computer Vision. Venice:IEEE,2017:2353-2362.
[4]
CaoZ, SimonT, WeiS H,et al.Realtime multi-person 2D pose estimation using part affinity fields[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition.Honolulu:IEEE,2017:1302-1310.
[5]
SunK, XiaoB, LiuD,et al.Deep high-resolution representation learning for human pose estimation[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Long Beach:IEEE,2019:5686-5696.
[6]
ChenY L, WangZ C, PengY X,et al.Cascaded pyramid network for multi-person pose estimation[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Salt Lake City:IEEE,2018:7103-7112.
[7]
LugaresiC, TangJ Q, NashH,et al. MediaPipe :a framework for building perception pipelines [EB/OL].(2019-06-14)[2024-03-11].
[8]
XiaR J, LiY S, LuoW H.LAGA-net:local-and-global attention network for skeleton based action recognition[J].IEEE Transactions on Multimedia,2021,24:2648-2661.
[9]
GaoX S, LiK Q, ZhangY,et al.3D skeleton-based video action recognition by graph convolution network[C]//2019 IEEE International Conference on Smart Internet of Things.Tianjin:IEEE,2019:500-501.
[10]
ShiL, ZhangY F, ChengJ,et al.Two-stream adaptive graph convolutional networks for skeleton-based action recognition[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Re-cognition.Long Beach:IEEE,2019:12018-12027.
[11]
YanS J, XiongY J, LinD H.Spatial temporal graph convolutional networks for skeleton-based action recognition[J].Proceedings of the AAAI Conference on Artificial Intelligence,2018,32(1):7444-7452.
[12]
SongL N, GuoX, FanY Q.Action recognition in video using human keypoint detection[C]//2020 15th International Conference on Computer Science & Education.Delft,Netherlands:IEEE,2020:465-470.
AlexiadisD S, DarasP.Quaternionic signal processing techniques for automatic evaluation of dance performances from MoCap data[J].IEEE Transactions on Multimedia,2014,16(5):1391-1406.