In the feature extraction of multimodal depression models, there are problems such as weak correlation between sentences, random feature fusion between different modalities, and lack of verification of the generalization ability of the model on the Chinese data set. By analyzing audio, text and visual features related to depression, this paper proposed a multi-modal depression recognition model STCMN(Sentence-level Temporal Convolutional Memory Network) based on improved TCN model. And the model was applied to the auxiliary diagnosis of clinical depression. Firstly, the fusion module of residual block, GRU and Self-Attention was used to extract the sentence-level features under different modalities, which enhances the context connection. Then, the TCN model was used to extract the global features of different modalities. Cross Attention was used to fuse the global features of different modalities mainly with multi-modal fusion features. Finally, the recognition results of the model for depression were obtained through the LogSoftmax layer. On the DAIC-WOZ public dataset, the accuracy rate, precision rate and recall rate of the proposed method for depression recognition reach 91.3%, 93.6% and 89.7%, respectively. The related indicators are better than other methods, which can better meet the needs of clinical medicine. On the private Chinese dataset MMD2022, the recognition results of STCMN model are still the best, indicating that the model has good generalization ability in Chinese depression recognition tasks.
SANTOMAUROD F, HERRERAA M M, SHADIDJ, et al. Global prevalence and burden of depressive and anxiety disorders in 204 countries and territories in 2020 due to the COVID-19 pandemic[J]. Lancet, 2021, 398(10312): 1700-1712.
[2]
MCGINNISE W, ANDERAUS P, HRUSCHAKJ, et al. Giving voice to vulnerable children: machine learning analysis of speech detects anxiety and depression in early childhood[J]. IEEE Journal of Biomedical and Health Informatics, 2019, 23(6): 2294-2301.
[3]
DI MATTEOD, FOTINOSK, LOKUGES, et al. The relationship between smartphone-recorded environmental audio and symptomatology of anxiety and depression: exploratory study[J]. JMIR Formative Research, 2020, 4(8): e18751.
[4]
FLORESR, TLACHACM L, TOTOE, et al. Transfer learning for depression screening from follow-up clinical interview questions[M].Singapore: Springer Nature Singapore, 2022: 53-78.
[5]
TLACHACM L, RUNDENSTEINERE. Screening for depression with retrospectively harvested private versus public text[J]. IEEE Journal of Biomedical and Health Informatics, 2020, 24(11): 3326-3332.
[6]
SENNS, TLACHACM L, FLORESR, et al. Ensembles of bert for depression classification[C]//2022 44th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC). IEEE, 2022: 4691-4694.
[7]
JAZAERY MAL, GUOG. Video-based depression level analysis by encoding deep spatiotemporal features[J]. IEEE Transactions on Affective Computing, 2021, 12(1): 262-268.
[8]
WANGQ, YANGH, YUY. Facial expression video analysis for depression detection in Chinese patients[J]. Journal of Visual Communication and Image Representation, 2018, 57: 228-233.
[9]
ASGARIM, SHAFRANI, SHEEBERL B. Inferring clinical depression from speech and spoken utterances[C]//2014 IEEE International Workshop on Machine Learning for Signal Processing(MLSP). IEEE, 2014: 1-5.
[10]
RODRIGUES MAKIUCHIM, WARNITAT, UTO K, et al. Multimodal fusion of bert-cnn and gated cnn representations for depression detection[C]//9th International Conference on Audio/Visual Emotion Challenge and Workshop, 2019: 55-63.
[11]
TOTOE, TLACHACM L, RundensteinerE A. Audibert: A deep transfer learning multimodal classification framework for depression screening[C]//30th ACM International Conference on Information & Knowledge Management, 2021: 4145-4154.
[12]
HAQUEA, GUOM, MINERA S, et al. Measuring depression symptom severity from spoken language and 3D facial expressions[DB/OL].(2018-11-21)[2023-09-27].
[13]
RAY A, KUMARS, REDDYR, et al. Multi-level attention network using text, audio and video for depression prediction[C]//9th International Conference on Audio/Visual Emotion Challenge and Workshop, 2019: 81-88.
[14]
CAOY, HAOY, LIB, et al. Depression prediction based on BiAttention-GRU[J]. Journal of Ambient Intelligence and Humanized Computing, 2022, 13(11): 5269-5277.
[15]
FLORESR, TLACHACM L, TOTOE, et al. AudiFace: Multimodal deep learning for depression screening[C]//Machine Learning for Healthcare Conference. PMLR, 2022: 609-630.
[16]
CHUNGJ, GULCEHREC, CHOK, et al. Gated feedback recurrent neural networks[C]//International Conference on Machine Learning, PMLR, 2015: 2067-2075.
[17]
WANGY Y, CHENJ, CHENX Q, et al. Short-term load forecasting for industrial customers based on TCN-LightGBM[J]. IEEE Transactions on Power Systems, 2021, 36(3): 1984-1997.
[18]
VIOLAP, JONESM J. Robust real-time face detection[J]. International Journal of Computer Vision, 2004, 57(2): 137-154.
[19]
MARTÍNEZ-CASTAÑOR, HTAITA, AZZOPARDIL, et al. Early risk detection of self-harm and depression severity using BERT-based transformers[C]//Proceedings of the Working Notes of CLEF,2020.
LIUHao, ZHUOGuangping, QIAOJunfu, et al. Chinese depression text classification based on domain emotion dictionary and word feature fusion[J]. Journal of North University of China (Natural Science Edition), 2022, 43(6): 522-529.(in Chinese)