In this work, an attention-based 3D convolutional neural network (ACN) is proposed for continuous sign language recognition in complex background. Firstly, the sign language video containing complex background is preprocessed with the background removal module. Then, the spatio-temporal fusion information is extracted by 3D-ResNet (3D residual convolutional neural network) based on spatial attention mechanism. Finally, the long short-term memory (LSTM) network combined with the time attention mechanism is used for sequence learning to obtain the final recognition result. Extensive experiments show that the algorithm performs well on the large-scale Chinese continuous sign language dataset CSL100. The algorithm shows good generalization performance facing different complex backgrounds, and the spatio-temporal attention mechanism introduced by the model is effective.
常见的手语视频往往连续数百帧,而对应的标签只有若干个汉字,二者在序列长度上相差巨大。为了解决这种长度差异带来的模型输入长度变化问题,学者们通常采取seq2seq(sequence to sequence)[25]的模型结构。另外,手语视频的帧与帧之间,存在大量的冗余信息,如何提取关键信息的同时剔除冗余信息显得尤为重要。为此,本文将注意力机制引入ACN网络。ACN网络采用编码器与解码器的结构,分别用来进行手语视频的时空特征提取和序列学习。
以第二个basic block为例,每个basic block中包含着1×1×1卷积核,3×3×3卷积核的三维卷积层以及batch normalize层,跳跃连接让浅层特征能直接传递至深层网络中,防止网络加深导致梯度爆炸的同时确保良好的性能。attention block对每帧的特征图进行3D全局最大池化(global max pooling,GMP)和3D全局平均池化(global average pooling,GAP),处理后得到两个特征映射,将其按照通道维数拼接在一起,得到特征映射。接着采用7×7×1三维卷积核的卷积层,对特征映射进行卷积,保证最终特征与输入特征映射在空间维度上的一致性。最后通过sigmoid激活函数得到最终特征。算法公式为
为了深入探究注意力机制的有效性,定量分析各注意力模块的贡献,本文进行注意力相关的消融实验。实验以结合LSTM的3D-ResNet18网络作为基本结构(baseline),分别考察空间注意力(spatial attention)模块、时间注意力(time attention)模块以及通道注意力(channel attention)模块对算法性能的影响,设计三种模型:1)B+S:baseline + spatial attention;2)B+S+T:baseline + spatial and time attention;3)B+S+T+C:baseline + spatial, time and channel attention。网络初始学习率为2E-3,批次大小为128(4块Tesla V100 GPU),采用Adam优化器进行优化训练,训练轮次设定为100轮,实验结果见图6与表4。
NIL, TANGW Y, HEZ Q, et al. Current situation and development trend of sign language service industry in China [J]. Language industry research, 2021,3:111-123(Ch).
MINAVALA, ALIFUK, XIEQ N, et al. Overview of sign language recognition methods and technologies [J]. Computer engineering and application, 2021,57 (18):1-12 (Ch).
[7]
王正胜,连淑红. 中国手语翻译研究二十年述评[J]. 译苑新谭,2021,2(1):99-108.
[8]
WANGZ S, LIANS H. A review of Chinese sign language translation in the past two decades [J]. New Perspectives in Translation Studies, 2021,2 (1):99-108 (Ch).
[9]
KAMALS M, CHENY D, LIS Z, et al. Technical approaches to Chinese sign language processing: A review[J]. IEEE Access, 2019,7:96926-96935. DOI:10.1109/ACCESS.2019.2929174 .
[10]
RASTGOOR, KIANIK, ESCALERAS. Sign language recognition: A deep survey [J]. Expert Systems with Applications, 2021,164:113794-113820. DOI:10.1016/j.eswa.2020.113794 .
[11]
ZHENGL, LIANGB. Sign language recognition using depth images [C]// Proceedings of International Conference on Control, Automation, Robotics and Vision(ICARCV). New York: IEEE Press, 2016:1-6. DOI:10.1109/ICARCV.2016.7838572 .
[12]
OLIVEIRAM, SUTHERLANDA, FAROUKM. Two-stage PCA with interpolated data for hand shape recognition in sign language [C]// Proceedings of IEEE Applied Imagery Pattern Recognition Workshop. New York: IEEE Press, 2016:1-4. DOI:10.1109/AIPR.2016.8010587 .
[13]
HASSANM, ASSALEHK, SHANABLEHT. User-dependent sign language recognition using motion detection[C]// Proceedings of International Conference on Computational Science and Computational Intelligence (CSCI). New York: IEEE Press, 2016:852-856. DOI:10.1109/CSCI.2016.0165 .
[14]
ZHANGJ H, ZHOUW G, LIH Q. A threshold-based HMM-DTW approach for continuous sign language recognition[C]// Proceedings of International Conference on Internet Multimedia Computing and Service. New York: ACM, 2014: 237-240. DOI:10.1145/2632856.2632931 .
[15]
VENUGOPALANS, ROHRBACHM, DONAHUEJ, et al. Sequence to sequence-video to text [C]// Proceedings of IEEE International Conference on Computer Vision. New York: IEEE Press, 2015:4534-4542. DOI:10.1109/ICCV.2015.515 .
[16]
YANGW W, TAOJ X, YEZ F. Continuous sign language recognition using level building based on fast hidden Markov model [J]. Pattern Recognition Letters, 2016,78:28-35. DOI:10.1016/j.patrec.2016.03.030 .
[17]
PIGOUL, DIELEMANS, KINDERMANSP J, et al. Sign language recognition using convolutional neural networks [C]// Proceedings of Workshop at the European Conference on Computer Vision. Cham: Springer International Publishing, 2014:1-6. DOI:10.1007/978-3-319-16178-5_40 .
[18]
KOLLERO, ZARGARANO, NEYH, et al. Deep sign: Hybrid CNN-HMM for continuous sign language recognition [C]// Proceedings of British Machine Vision Conference. York: British Machine Vision Association, 2016:1-12. DOI:10.5244/c.30.136 .
[19]
XIAOQ K, ZHAOY D, HUANW. Multi-sensor data fusion for sign language recognition based on dynamic Bayesian network and convolutional neural network [J]. Multimedia Tools and Applications, 2019, 78(11):15335-15352. DOI:10.1007/s11042-018-6939-8 .
[20]
HUANGJ, ZHOUW G, LIH Q, et al. Sign language recognition using 3D convolutional neural networks [C]// Proceedings of International Conference on Multimedia and Expo(ICME). New York: IEEE Press, 2015:1-6. DOI:10.1109/ICME.2015.7177428 .
[21]
SONGP P, GUOD, XINH R, et al. Parallel temporal encoder for sign language translation [C]// Proceedings of the IEEE International Conference on Image Processing (ICIP). New York: IEEE Press, 2019:1915-1919. DOI:10.1109/ICIP.2019.8803123 .
[22]
ZHOUM J, NG M, CAIZ X, et al. Self-attention-based fully-inception networks for continuous sign language recognition [OL/DB]. [2021-08-18].
[23]
PUJ, ZHOUW, LIH. Sign language recognition with multimodal features[C]// Proceedings of Pacific Rim Conference on Multimedia. Switzerland: Springer, 2016:252-261. DOI: 10.1007/978-3-319-48896-7_25 .
[24]
QINW Y, MEIX, CHENY M, et al. Sign language recognition and translation method based on VTN [C]// Proceedings of 2021 International Conference on Digital Society and Intelligent Systems (DSInS). New York: IEEE Press, 2021:111-115. DOI:10.1109/DSInS54396.2021.9670588 .
[25]
ZOUR L, MAK B. C2SLR: Consistency-enhanced continuous sign language recognition[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2022:5131-5140. DOI:10.1109/CVPR.2019.00429 .
[26]
LINS C, RYABTSEVA, SENGUPTAS, et al. Real-time high-resolution background matting [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2021:8762-8771. DOI:10.1109/CVPR46437.2021.00865 .
[27]
CHENL C, PAPANDREOUG, KOKKINOSI, et al. DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, 40(4):834-848. DOI:10.1109/TPAMI.2017.2699184 .
[28]
HINTONG E, SALAKHUTDINOVR R. Reducing the dimensionality of data with neural networks[J]. Science, 2006, 313(5786):504-507. DOI:10.1126/science.1127647 .
[29]
SUTSKEVERI, VINYALSO, LEQ. Sequence to sequence learning with neural networks [C]// Proceedings of the 27th International Conference on Neural Information Processing Systems (NIPS'14). New York: ACM, 2014(2):3104-3112.
[30]
WILLIAMSR J, ZIPSERD. A learning algorithm for continually running fully recurrent neural networks[J] Neural Computation, 1989,1(2):270-280. DOI:10.1162/neco.1989.1.2.270 .
[31]
DREUWP, NEIDLEC, ATHITSOSV, et al. Benchmark databases for video-based automatic sign language recoginition[C]// Proceedings of the 6th International Conference on Language Resources and Evaluation. Paris: ELRA, 2008:1-6.
[32]
ADALOGLOUN, CHATZIST, PAPASTRATISI, et al. A comprehensive study on deep learning-based methods for sign language recognition [EB/OL].2020:arXiv:2007.12530.
[33]
FORSTERJ, SCHMIDTC, HOYOUXT, et al. RWTH-PHOENIX-Weather: A large vocabulary sign language recognition and translation corpus [C]// Proceedings of the 8th International Conference on Language Resources and Evaluation. Paris: ELRA, 2012:3785-3789.
[34]
CAMGOZN C, HADFIELDS, KOLLERO, et al. Neural sign language translation [C]// 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018:7784-7793. DOI:10.1109/CVPR.2018.00812 .
[35]
HUANGJ, ZHOUW G, ZHANGQ L, et al. Video-based sign language recognition without temporal segmentation[C]// Proceedings of AAAI Conference on Artificial Intelligence. Menlo Park: AAAI Press, 2018:2257-2264. DOI:10.1609/aaai.v32i1.11903 .
[36]
DUARTEA C. Cross-modal neural sign language translation [C]// Proceedings of the 27th ACM International Conference on Multimedia. New York: ACM, 2019:1650-1654. DOI:10.1145/3343031.3352587 .
[37]
CAMGOZN C, HADFIELDS, KOLLERO, et al. SubUNets: End-to-end hand shape and continuous sign language recognition [C]// 2017 IEEE International Conference on Computer Vision. New York: IEEE Press, 2017:3075-3084. DOI:10.1109/ICCV.2017.332 .