Stuttering speech classification aims to classify and recognize different categories of stuttering using spoken signals. Nevertheless, the existing related works fail to sufficiently focus on sequential characteristics for the representation embedding of self-supervised pre-trained models, and these works also simplistically address the class-imbalance issue for stuttering-speech data. In this regard, we proposed a stuttering speech classification approach based on self-supervised pre-trained models and nonlinear weighted cross-entropy (NWCE) loss. Within the proposed approach, we first employed a self-supervised pre-trained model to extract paralinguistic representation embeddings from stuttering speech. Then, we utilized a bidirectional long short-term memory network model with a self-attention mechanism to capture essential temporal features and contextual information within the embeddings. Afterwards, a nonlinear weighted cross-entropy loss was performed to focus on stuttering speech categories with fewer samples. The experimental results on stuttering speech classification dataset indicate that, the proposed approach achieves better performance for classifying stuttering speech compared with state-of-the-art approaches, through learning the sequential information from self-supervised pre-trained models’ multi-layer representation embedding in speech, and sufficiently describes the relationship between the data of different stuttering categories by using NWCE.
早期相关研究主要基于声学特征的提取和分析,并采用了深度学习基础模型。Sheikh等[6]利用梅尔频率倒谱系数(Mel-scale Frequency Cepstral Coefficients,MFCC)和时延神经网络(Time Delay Neural Network,TDNN)描述口吃语音的时序信息。Kourkunakis等[7]则使用频谱图作为输入,结合深度残差网络(Residual Networks,ResNets)和双向长短期记忆网络(Bidirectional Long Short Term Memory Networks,Bi-LSTM),对口吃语音进行了多分类。然而,由于口吃病例数量有限、个人隐私的保护以及获取标准化数据的难度较大等因素,口吃数据样本稀少[8],导致无法很好地捕捉口吃语音中的副语言信息[9]。于是,Grósz等[10]将自监督预训练模型应用于口吃语音分类,得到了更通用的副语言表示嵌入。
SCHULLERB W, BATLINERA, AMIRIPARIANS,et al.The ACM multimedia 2023 computational paralinguistics challenge:Emotion share & requests[C]//ACM International Conference on Multimedia,2023:9635-9639.
[2]
XUX, DENGJ, ZHANGZ,et al.Zero-shot speech emotion recognition using generative learning with reconstructed prototypes[C]//IEEE International Conference on Acoustics,Speech and Signal Processing (ICASSP).IEEE,2023:1-5.
[3]
HAIDERF, DE LA FUENTES,LUZ S.An assessment of paralinguistic acoustic features for detection of Alzheimer's dementia in spontaneous speech[J].IEEE Journal of Selected Topics in Signal Processing,2019,14(2):272-281.
[4]
SHAHINM, ZAFARU, AHMEDB.The automatic detection of speech disorders in children:Challenges,opportunities,and preliminary results[J].IEEE Journal of Selected Topics in Signal Processing,2019,14(2):400-412.
[5]
SHEIKHS A, SAHIDULLAHM, HIRSCHF,et al.Machine learning for stuttering identification:Review,challenges and future directions[J].Neurocomputing,2022,514:385-402.
[6]
SHEIKHS A, SAHIDULLAHM, HIRSCHF,et al.Stutternet:Stuttering detection using time delay neural network[C]//European Signal Processing Conference (EUSIPCO).IEEE,2021:426-430.
[7]
KOURKOUNAKIST, HAJAVIA, ETEMADA.Detecting multiple speech disfluencies using a deep residual network with bidirectional long short-term memory[C]//IEEE International Conference on Acoustics,Speech and Signal Processing(ICASSP).IEEE,2020:6089-6093.
[8]
SEBASTIANP B, DOMINIKW, ELMARN,et al.Detecting dysfluencies in stuttering therapy using wav2vec 2.0[C]//Annual Conference of the International Speech Communication Association (INTERSPEECH),2022: 347.
[9]
BAYERLS P, WAGNERD, NÖTHE,et al.The influence of dataset partitioning on dysfluency detection systems[C]//International Conference on Text,Speech,and Dialogue(ICTSD).Cham:Springer International Publishing,2022:423-436.
[10]
GRÓSZT, PORJAZOVSKID, GETMANY,et al.wav2vec2-based paralinguistic systems to recognise vocalised emotions and stuttering[C]//ACM International Conference on Multimedia,2022:7026-7029.
[11]
SHEIKHS A, SAHIDULLAHM, OUNIS,et al.End-to-end and self-supervised learning for ComParE 2022 stuttering sub-challenge[C]//ACM International Conference on Multimedia,2022:7104-7108.
[12]
JOUAITIM, DAUTENHAHNK.Dysfluency classification in stuttered speech using deep learning for real-time applications[C]//IEEE International Conference on Acoustics,Speech and Signal Processing (ICASSP).IEEE,2022:6482-6486.
[13]
KOURKOUNAKIST, HAJAVIA, ETEMADA.FluentNet:End-to-end detection of stuttered speech disfluencies with deep learning[J].IEEE/ACM Transactions on Audio,Speech,and Language Processing,2021,29:2986-2999.
[14]
BAEVSKIA, ZHOUY, MOHAMEDA,et al.wav2vec 2.0:A framework for self-supervised learning of speech representations[J].Advances in Neural Information Processing Systems,2020,33:12449-12460.
[15]
SUNH, LIANZ, LIUB,et al.EmotionNAS:Two-stream architecture search for speech emotion recognition[DB/OL].(2022-03-25)[2023-09-05].
[16]
VAESSENN, VAN LEEUWEND A.Fine-tuning wav2vec2 for speaker recognition[C]//IEEE International Conference on Acoustics,Speech and Signal Processing (ICASSP).IEEE,2022:7967-7971.
[17]
BRAUNF, ERZIGKEITA, LEHFELDH,et al.Going beyond the cookie theft picture test:Detecting cognitive impairments using acoustic features[C]//International Conference on Text,Speech,and Dialogue (ICTSD).Cham:Springer International Publishing,2022:437-448.
[18]
GHOSHS, TYAGIU, KUMARS,et al.A novel multimodal dynamic fusion network for disfluency detection in spoken utterances[DB/OL].(2022-11-27)[2023-09-05].
[19]
SHARMAM.Multi-lingual multi-task speech emotion recognition using wav2vec 2.0[C]//IEEE International Conference on Acoustics,Speech and Signal Processing (ICASSP).IEEE,2022:6907-6911.
[20]
RENZ, NGUYENT T, CHANGY,et al.Fast yet effective speech emotion recognition with self-distillation[C]//IEEE International Conference on Acoustics,Speech and Signal Processing (ICASSP).IEEE,2023:1-5.
[21]
BAYERLS P, WAGNERD, BAUMANNI,et al.Detecting vocal fatigue with neural embeddings[DB/OL].(2022-04-07)[2023-09-05].
[22]
BAYERLS P, GERCZUKM, BATLINERA,et al.Classification of stuttering-The ComParE challenge and beyond[J].Computer Speech & Language,2023,81:101519.
[23]
MONTACIÉC, CARATYM J, LACKOVICN.Audio features from the wav2vec 2.0 embeddings for the ACM multimedia 2022 stuttering challenge[C]//ACM International Conference on Multimedia,2022:7195-7199.
[24]
SCHULLERB, BATLINERA, AMIRIPARIANS,et al.The ACM multimedia 2022 computational paralinguistics challenge:Vocalisations,stuttering,activity,& mosquitoes[C]//ACM International Conference on Multimedia,2022:7120-7124.
[25]
BAYERLS, VON GUDENBERGA W, HÖNIGF,et al.KSoF:The Kassel State of Fluency dataset-a therapy centered dataset of stuttering[C]//Language Resources and Evaluation Conference (LREC),2022:1780-1787.