面向精准育人的多模态学生言语情感识别方法研究

刘振中 ,  张夷楠 ,  王翠萍

河南师范大学学报(自然科学版) ›› 2026, Vol. 54 ›› Issue (4) : 98 -105.

PDF (5787KB)
河南师范大学学报(自然科学版) ›› 2026, Vol. 54 ›› Issue (4) : 98 -105. DOI: 10.16366/j.cnki.1000-2367.2025.10.30.0002
数学与计算机科学

面向精准育人的多模态学生言语情感识别方法研究

作者信息 +

Research on multimodal student speech emotion recognition methods for precise education

Author information +
文章历史 +
PDF (5925K)

摘要

本文面向国家“精准育人”战略需求,针对课堂场景中学生面部表情变化细微,导致现有基于计算机视觉的情感识别方法面临识别困难的现实瓶颈,创新性地从学生言语维度出发,提出了一种融合多维度声学特征与文本特征的多模态学生言语情感识别方法,并构建了多模态感知神经网络.其中,在该网络的声学编码器中设计了注意力统计池化层,能够有效提取学生言语中的关键帧级别特征和序列级特征.实验结果表明,所提出的方法中的不同模块均能有效提升情感识别性能,且在识别效果上显著优于现有单模态方法.该方法有效融合了学生言语中的声学和文本的信息,突破了单模态感知的局限性,能够为“精准育人”实践提供可靠的技术支持,切实提高育人成效.

Abstract

This study addressed the national strategic requirement of precision education by focusing on the subtle changes in students' facial expressions in classroom, which pose significant challenges to existing computer vision-based emotion recognition methods. This study innovatively approach the issue from the dimension of student speech, proposing a multimodal student speech emotion recognition method that integrates multidimensional acoustic and textual features, and constructing a multimodal perception neural network. Specifically, an attention statistical pooling layer is designed in the acoustic encoder of this network, effectively extracting key frame-level and sequence-level features from student speech. Experimental results demonstrate that the different modules within the proposed method significantly enhance emotion recognition performance, outperforming existing unimodal methods in recognition effectiveness. This method effectively integrates acoustic and textual information from student speech, overcoming the limitations of unimodal perception, and provides reliable technical support for the practice of precision education, thereby substantially improving educational outcomes.

关键词

精准育人 / 学生言语情感 / 声学特征 / 多模态情感识别

Key words

precise education / students' speech emotion / acoustic features / multimodal emotion recognition

引用本文

引用格式 ▾
刘振中,张夷楠,王翠萍. 面向精准育人的多模态学生言语情感识别方法研究[J]. 河南师范大学学报(自然科学版), 2026, 54(4): 98-105 DOI:10.16366/j.cnki.1000-2367.2025.10.30.0002

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

彭婧, 陈强. 从 “选育”到 “孕育”:中美基础学科拔尖创新人才培养模式比较与启示[J]. 创新科技, 2025, 25(11): 11-22.

[2]

Peng J, Chen Q. From "selection and cultivation" to "nurturing": a comparative study of training models for innovative talents in fundamental disciplines between China and the United States[J]. Innovation Science and Technology, 2025, 25(11): 11-22.

[3]

方海光, 洪心, 孔新梅, 等 . 基于课堂交互行为数据的教学教研融合研究:以网络画板和iFIAS的数学教学教研应用为例[J]. 中国电化教育, 2022(5): 115-121.

[4]

Fang H G, Hong X, Kong X M, et al. Research on teaching and research integration based on classroom interactive behavior data: a case application of net pad and iFIAS in mathematics[J]. China Educational Technology, 2022(5): 115-121.

[5]

李协吉, 孙洪涛, 高佩琪. 教师-学生积极情感表达一致性对中学生学习投入的影响:基于多项式回归与响应面分析[J]. 教师教育研究, 2024, 36(4): 72-80.

[6]

Li X J, Sun H T, Gao P Q. The influence of teacher-students' Positive emotional expression consistency on middle school Students' Learning engagement: based on polynomial regression and response surface analysis[J]. Teacher Education Research, 2024, 36(4): 72-80.

[7]

张屹, 郝琪, 陈蓓蕾, 等 . 智慧教室环境下大学生课堂学习投入度及影响因素研究:以“教育技术学研究方法课”为例[J]. 中国电化教育, 2019(1): 106-115.

[8]

Zhang Y, Hao Q, Chen B L, et al. Research on college students' classroom engagement and its influencing factors in smart classroom environment: using educational technology research method course as an example[J]. China Educational Technology, 2019(1): 106-115.

[9]

Dukic D, Sovic Krzic A . Real-time facial expression recognition using deep learning with application in the active classroom environment[J]. Electronics, 2022, 11(8): 1240.

[10]

Cámbara G, Luque J, Farrús M . Convolutional speech recognition with pitch and voice quality features[EB/OL]. [2025-09-12]. https://arxiv.org/abs/2009.01309.

[11]

黄庭培, 郑秋梅, 李世宝. 教师和学生的课堂行为互动及优化策略[J]. 教育理论与实践, 2016, 36(32): 54-56.

[12]

Huang T P, Zheng Q M, Li S B. Interaction between teachers and students' classroom behavior and optimization strategies[J]. Theory and Practice of Education, 2016, 36(32): 54-56.

[13]

杨伊, 陈昌来, 陈兴冶. 基于多模态语料库的教师评价类话语特征研究[J]. 教师教育研究, 2024, 36(5): 22-29.

[14]

Yang Y, Chen C L, Chen X Y. A study on the characteristics of teacher evaluation discourse based on multimodal corpus[J]. Teacher Education Research, 2024, 36(5): 22-29.

[15]

郭子漾, 李国勇. 基于双门限的语音端点检测算法改进[J]. 计算机应用, 2025, 45(S1): 101-105.

[16]

Guo Z Y, Li G Y. Improvement of speech endpoint detection algorithm based on dual-threshold[J]. Journal of Computer Applications, 2025, 45(S1): 101-105.

[17]

Wang C Y, Ren Y, Zhang N, et al. Speech emotion recognition based on multi-feature and multi-lingual fusion[J]. Multimedia Tools and Applications, 2022, 81(4): 4897-4907.

[18]

Gao Z F, Zhang S L, McLoughlin I, et al. Paraformer: fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition[EB/OL]. [2025-09-12]. https://arxiv.org/abs/2206.08317.

[19]

Devlin J, Chang M W, Lee K, et al. BERT: pre-training of deep bidirectional transformers for language understanding[C]// Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (long and Short Papers). Minneapolis: [s. n.], 2019: 4171-4186.

[20]

Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[EB/OL]. [2025-09-13]. https://proceedings.neurips.cc/paper_files/paper/2017/hash3f5ee243547dee91fbd053c1c4a845aa-Abstract.html.

[21]

孙志, 王冠. 自监督对比学习的CNN-GRU语音情感识别算法[J]. 西安电子科技大学学报(自然科学版), 2024, 51(6): 182-193.

[22]

Sun Z, Wang G. CNN-GRU speech emotion recognition algorithm for self-supervised comparative learning[J]. Journal of Xidian University (Natural Science), 2024, 51(6): 182-193.

[23]

Busso C, Bulut M, Lee C C, et al. IEMOCAP: interactive emotional dyadic motion capture database[J]. Language Resources and Evaluation, 2008, 42(4): 335-359.

[24]

Poria S, Hazarika D, Majumder N, et al. MELD: a multimodal multi-party dataset for emotion recognition in conversations[EB/OL]. [2025-09-13]. https://arxiv.org/abs/1810.02508.

[25]

Yi Y F, Tian Y, He C, et al. DBT: multimodal emotion recognition based on dual-branch transformer[J]. The Journal of Supercomputing, 2023, 79(8): 8611-8633.

[26]

Majumder N, Hazarika D, Gelbukh A, et al. Multimodal sentiment analysis using hierarchical fusion with context modeling[J]. Knowledge-Based Systems, 2018, 161: 124-133.

[27]

Makiuchi M R, Uto K, Shinoda K . Multimodal emotion recognition with high-level speech and text features[C]// 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). Cartagena: IEEE, 2021: 350-357.

[28]

Santoso J, Yamada T, Ishizuka K, et al. Speech emotion recognition based on self-attention weight correction for acoustic and text features[J]. IEEE Access, 2022, 10: 115732-115743.

[29]

Poria S, Cambria E, Hazarika D, et al. Context-dependent sentiment analysis in user-generated videos[C]// Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (volume 1: Long Papers). [S. l.]: ACL, 2017: 873-883.

[30]

Ghosal D, Majumder N, Poria S, et al. DialogueGCN: a graph convolutional neural network for emotion recognition in conversation[EB/OL]. [2025-09-16]. https://arxiv.org/abs/1908.11540.

[31]

Hu D, Wei L W, Huai X Y . DialogueCRN: contextual reasoning networks for emotion recognition in conversations[EB/OL]. [2025-09-13]. https://arxiv.org/abs/2106.01978.

[32]

Hu D, Hou X L, Wei L W, et al. MM-DFN: multimodal dynamic fusion network for emotion recognition in conversations[C]// ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Singapore: IEEE, 2022: 7037-7041.

基金资助

国家社会科学基金教育学青年项目(CBA230291)

河南省软科学研究计划项目(242400410444)

AI Summary AI Mindmap
PDF (5787KB)

0

访问

0

被引

详细

导航
相关文章

AI思维导图

/

〈 〉