2.Faculty of Education,The University of Hong Kong,Hong Kong 999077,China
3.Faculty of Artificial Intelligence in Education,Central China Normal University,Wuhan 430074,Hubei,China
Show less
文章历史+
Received
Published
2025-01-19
2025-10-24
Issue Date
2026-07-23
PDF (9185K)
摘要
面部表情识别(Facial Expression Recognition,FER)是计算机视觉领域的重要研究方向,但在教育场景的实际应用中仍面临诸多挑战,如遮挡、光照变化和头部姿态干扰等问题。本文定义了褶纹基元和褶纹基元关系;采用褶纹基元刻画面部局部区域因肌肉形变引发的方向性纹理特征,利用褶纹基元关系刻画其几何关联;指出褶纹基元具有高斯亲和、类间共享和类内不变特征。在此基础上,提出一种基于褶纹基元感知的高效面部表情识别模型P2FER(Pleated Primitive for Facial Expression Recognition)。在该模型中,设计了高斯亲和褶纹基元提取模块,用于捕获面部表情特征的多样性。通过引入高斯亲和褶纹基元,模型能够学习更具区分性的特征,增强特征空间的多样性,从而对不同类别表情的分类起到关键作用。为进一步增强模型对不同表情特征的建模能力,本文提出褶纹基元线索感知模块。该模块通过构建正负样本对,结合类关系设计了一种新颖的损失函数,实现类间共享褶纹基元与类内不变褶纹基元的协同挖掘,借助于不同类别间褶纹基元关联关系建模。该模块显著提升了模型对表情的识别能力。在RAF-DB和FERPlus数据集上的实验表明,本文模型在多项指标上均优于现有方法,性能提升显著。实验结果验证了高斯亲和褶纹基元提取模块与褶纹基元线索感知模块的有效性,表明用褶纹基元关系建模可增强模型的泛化能力。
Abstract
Facial Expression Recognition (FER) is an important research direction in the field of computer vision. However, it still faces many challenges in practical applications in educational scenarios, such as occlusion, illumination changes, and interference from head postures. This paper defines pleated primitives and pleated primitive relationships. Pleated primitives are used to characterize directional texture features in local facial regions induced by muscle deformations, while pleated primitive relationships are employed to depict their geometric correlations. It is indicated that pleated primitives possess Gaussian affinity, inter-class sharing,and intra-class invariance characteristics. On this basis, an efficient facial expression recognition model P2FER(Pleated Primitive for Facial Expression Recognition) based on pleated primitive perception is proposed. In this model, a Gaussian affinity pleated primitive extractor was designed to capture the diversity of facial expression features. By introducing Gaussian affinity pleated units, the model can learn more discriminative features and enhance the diversity of the feature space, thereby playing a key role in the classification of different types of expressions. To further enhance the modeling ability of the model for different expression features, this paper proposes a pleated primitive clue perception model. This model constructs positive and negative sample pairs and designs a novel loss function in combination with class relationships to achieve the collaborative mining of shared pleated primitives between classes and invariant pleated primitives within classes. With the modeling of the correlation relationship of pleated primitives among different categories, this module significantly improves the model's ability to recognize expressions. Experiments on the RAF-DB and FERPlus datasets show that the model proposed in this paper outperforms existing methods in multiple indicators and has significant performance improvement. The experimental results verified the effectiveness of the Gaussian affinity pleated primitive extractor and the pleated primitive cue perception module and proved that modeling using the pleated primitive relationship can enhance the generalization ability of the model.
基于CNN的FER方法:2019年,Wang等[4]开发的LS(Local and mutil-Scale)-CNN提出将多尺度特征提取、空间注意和通道注意相结合改进人脸识别,在各种人脸识别任务中表现优异。Li等[5]开发了一种新颖的带有注意力模型的卷积神经网络,用于面部表情识别,该网络利用局部二元模式(LBP)提取图像纹理信息,捕捉面部细微动作,提高网络性能。Liu等[6]提出BPMB(Bayesian Convolutional Neural Network with Perturbal Mutil-Branch Structrue)模型,通过结合变分推断和贝叶斯反向传播,有效量化了人脸表情识别中的不确定性,并通过引入扰动多分支结构和Dropout机制,在训练过程中增强了模型的鲁棒性,提升了深层特征的提取能力,从而在噪声数据和不确定性环境下实现了更稳健的表现。
JIANGC S, LIUZ T, WUM, et al. Efficient facial expression recognition with representation reinforcement network and transfer self-training for human-machine interaction[J]. IEEE Transactions on Industrial Informatics, 2023, 19(9): 9943-9952. DOI: 10.1109/TII.2022.3233650 .
[2]
BISOGNIC, CASTIGLIONEA, HOSSAINS, et al. Impact of deep learning approaches on facial expression recognition in healthcare industries[J]. IEEE Transactions on Industrial Informatics, 2022, 18(8): 5619-5627. DOI: 10.1109/TII.2022.3141400 .
XUY P, LIY M, GUOW T, et al. Intelligent assessment of classroom learning engagement integrating behavior and emotion analysis [J/OL]. J Wuhan Univ (Nat Sci Ed), 2025: 1-13. DOI:10.14188/j.1671-8836.2024.0164(Ch ).
[5]
WANGQ C, GUOG D. LS-CNN: Characterizing local patches at multiple scales for face recognition[J]. IEEE Transactions on Information Forensics and Security, 2019, 15: 1640-1653. DOI: 10.1109/TIFS.2019.2946938 .
[6]
LIJ, JINK, ZHOUD L, et al. Attention mechanism-based CNN for facial expression recognition[J]. Neurocomputing, 2020, 411: 340-350. DOI: 10.1016/j.neucom.2020.06.014 .
[7]
LIUS S, ZHAOD X, SUNZ B, et al. BPMB: BayesCNNs with perturbed multi-branch structure for robust facial expression recognition[J]. Image and Vision Computing, 2024, 143: 104960. DOI: 10.1016/j.imavis.2024.104960 .
[8]
XUEF L, WANGQ C, GUOG D. TransFER: Learning relation-aware facial expression representations with transformers[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE Press, 2021: 3581-3590. DOI: 10.1109/ICCV48922.2021.00358 .
[9]
LIH T, SUIM Z, ZHAOF, et al. MVT: Mask vision transformer for facial expression recognition in the wild [EB/OL]. [2025-01-02].
[10]
LEEI, LEEE, YOOS B. Latent-OFER: Detect, mask, and reconstruct with latent vectors for occluded facial expression recognition[C]//2023 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE Press, 2023: 1536-1546. DOI: 10.1109/ICCV51070.2023.00148 .
[11]
YEJ Y, YUY H, ZHENGY S, et al. Dep-FER: Facial expression recognition in depressed patients based on voluntary facial expression mimicry[J]. IEEE Transactions on Affective Computing, 2024, 15(3): 1725-1738. DOI: 10.1109/TAFFC.2024.3370103 .
[12]
MAF Y, SUNB, LIS T. Transformer-augmented network with online label correction for facial expression recognition[J]. IEEE Transactions on Affective Computing, 2024, 15(2): 593-605. DOI: 10.1109/TAFFC.2023.3285231 .
[13]
LIS, DENGW H, DUJ P. Reliable crowdsourcing and deep locality-preserving learning for expression recognition in the wild[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2017: 2584-2593. DOI: 10.1109/CVPR.2017.277 .
[14]
BARSOUME, ZHANGC, FERRERC C, et al. Training deep networks for facial expression recognition with crowd-sourced label distribution[C]//Proceedings of the 18th ACM International Conference on Multimodal Interaction. New York: ACM, 2016: 279-283. DOI: 10.1145/2993148.2993165 .
[15]
ZHAOS W, CAIH B, LIUH H, et al. Feature selection mechanism in CNNs for facial expression recognition [EB/OL]. [2025-02-03].
[16]
KUOC M, LAIS H, SARKISM. A compact deep learning model for robust facial expression recognition[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). New York: IEEE Press, 2018: 2202-22028. DOI: 10.1109/CVPRW.2018.00286 .
[17]
FANY R, LAMJ C K, LIV O K. Multi-region ensemble convolutional neural network for facial expression recognition[M]//Artificial Neural Networks and Machine Learning ― ICANN 2018. Cham: Springer International Publishing, 2018: 84-94. DOI: 10.1007/978-3-030-01418-6_9 .
[18]
LIH Y, WANGN N, YUY, et al. LBAN-IL: A novel method of high discriminative representation for facial expression recognition[J]. Neurocomputing, 2021, 432: 159-169. DOI: 10.1016/j.neucom.2020.12.076 .
[19]
PUT, CHENT S, XIEY, et al. AU-expression knowledge constrained representation learning for facial expression recognition[C]//2021 IEEE International Conference on Robotics and Automation (ICRA). New York: IEEE Press, 2021: 11154-11161. DOI: 10.1109/icra48506.2021.9561252 .
[20]
HUANGQ H, HUANGC Q, WANGX Z, et al. Facial expression recognition with grid-wise attention and visual transformer[J]. Information Sciences, 2021, 580: 35-54. DOI: 10.1016/j.ins.2021.08.043 .
[21]
FARZANEHA H, QIX J. Facial expression recognition in the wild via deep attentive center loss[C]//2021 IEEE Winter Conference on Applications of Computer Vision (WACV). New York: IEEE, 2021: 2401-2410. DOI: 10.1109/wacv48630.2021.00245 .
[22]
LIY J, LUG M, LIJ X, et al. Facial expression recognition in the wild using multi-level features and attention mechanisms[J]. IEEE Transactions on Affective Computing, 2023, 14(1): 451-462. DOI: 10.1109/TAFFC.2020.3031602 .
[23]
WANGK, PENGX J, YANGJ F, et al. Suppressing uncertainties for large-scale facial expression recognition[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2020: 6896-6905. DOI: 10.1109/CVPR42600.2020.00693 .
[24]
ZHOUB Y, CUIQ, WEIX S, et al. BBN: Bilateral-branch network with cumulative learning for long-tailed visual recognition[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE, 2020: 9716-9725. DOI: 10.1109/cvpr42600.2020.00974 .
[25]
ZHANGY H, LIY Q, QINL X, et al. Leave no stone unturned: Mine extra knowledge for imbalanced facial expression recognition [EB/OL]. [2025-03-01]. DOI: 10.1109/iciecs.2009.5362811 .
[26]
CHENX C, ZHENGX W, SUNK, et al. Self-supervised vision transformer-based few-shot learning for facial expression recognition[J]. Information Sciences, 2023, 634: 206-226. DOI: 10.1016/j.ins.2023.03.105 .
[27]
CHEND L, WENG H, LIH H, et al. Multi-relations aware network for in-the-wild facial expression recognition[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2023, 33(8): 3848-3859. DOI: 10.1109/TCSVT.2023.3234312 .
[28]
MAF Y, SUNB, LIS T. Facial expression recognition with visual transformers and attentional selective fusion[EB/OL]. 2021: 2103.16854. DOI: 10.1109/taffc.2021.3122146 .
[29]
ZHANGY H, WANGC R, LINGX, et al. Learn from all: Erasing attention consistency for noisy label facial expression recognition[C]//Computer Vision ― ECCV 2022. Cham: Springer, 2022: 418-434. DOI: 10.1007/978-3-031-19809-0_24 .
[30]
ZHENGC, MENDIETAM, CHENC. POSTER: A pyramid cross-fusion transformer network for facial expression recognition[C]//2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW). New York: IEEE Press, 2023: 3138-3147. DOI: 10.1109/ICCVW60793.2023.00339 .
[31]
MAATENL, DERVAN, HINTONG. Visualizing data using t-SNE [J]. Journal of Machine Learning Research, 2008, 9(86): 2579-2605.
[32]
SELVARAJUR R, COGSWELLM, DAS A, et al. Grad-CAM: Visual explanations from deep networks via gradient-based localization[J]. International Journal of Computer Vision, 2020, 128(2): 336-359. DOI: 10.1007/s11263-019-01228-7 .