Tibetan text recognition holds significant application value in rapid text input, digital preservation, and smart city development. To advance the research and development of Tibetan text recognition technologies, this paper presents a comprehensive review and analysis of recognition algorithms for Tibetan script. First, it introduces the fundamental concepts of Tibetan text recognition, including definitions and distinctions among various recognition types, along with relevant research background and primary research pathways. Based on different scripts and writing styles, Tibetan text recognition is categorized into four types: modern printed Tibetan recognition, handwritten Tibetan recognition, Tibetan scene text recognition, and Tibetan ancient document recognition. Next, the paper elaborates on the application of traditional OCR methods and deep learning techniques to Tibetan text recognition, highlighting the research shift toward deep learning approaches, which has significantly improved performance across various tasks. The paper also provides a brief overview of existing datasets used for Tibetan text recognition. Finally, it discusses the current limitations in different recognition scenarios and outlines potential future research directions.
为进一步提升识别性能,近年来有研究人员提出了多种文本行切分方法[65-71]。李金成等[72]在以上研究基础上,提出了一种基于文本核心区域扩展生长的文本行分割方法,该方法通过提取藏文古籍文档图像中的文本核心区域,在其基础上扩展生长,以实现文本行的有效切分。此外,贡去卓么等[73]针对藏文古籍文档的文本区域检测问题,提出了一种基于图像语义分割的方法。该方法首先利用判别式对抗网络框架下的语义分割网络,对藏文古籍文档中不同类型的文本区域进行像素级分类;随后,根据分类结果提取各个文本区域的轮廓;最后,将检测到的版面布局信息进行保存。实验结果显示,该方法在藏文木刻版印刷体古籍文献的文本区域检测任务中取得了较为理想的效果。在此基础上,仁青东主等[59]针对藏文古籍木刻版文献在复杂图像质量、文字粘连等方面的识别难题,提出了一种CNN与RNN融合的神经网络识别模型。该模型采用基于滑动窗的行识别技术,解决了行文字较长导致的粘连问题。同时,采用串识别技术构建串识别框架,其核心基于残差网络(Residual Network, ResNet),结合连接时域分类(Connectionist Temporal Classification, CTC)和双向长短期记忆网络(Bidirectional Long Short-Term Memory, Bi-LSTM)完成块图像的识别(图4)。此外,为进一步提升模型的泛化能力,该研究引入了合成样本[74]进行预训练,有效提升了模型在藏文古籍文字识别中的性能表现。
藏文古籍手写本中大多采用乌梅体字体。乌梅体因其书写形式多样、文字模糊、笔画弯曲、字迹褪色、字符残缺、断裂粘连、字体大小不一、字体风格差异较大、文本行间距不均等特点,显著增加了手写体藏文古籍文献的文字识别难度。因此,研究文献较少,相关技术发展相对滞后。2020年王梦锦等[71]对过去十余年自然场景文本检测领域中常用的算法及其研究趋势进行了梳理,并采用了CTPN(Connectionist Text Proposal Network)算法与EAST(Efficient and Accurate Scene Text Detector)对手写体藏文古籍文本进行检测。然而,该检测方法的数据集标注方式为矩形框,且采用了基于回归框的检测机制,而藏文古籍文本行间距较小,且藏文字存在上加字与下加字的结构特点,易导致检测框之间出现大量重叠,进而影响最终的手写体藏文古籍文本检测性能。为进一步提升手写体藏文古籍文本的检测效果,芷香香等[75]对手写体藏文古籍文本进行了深入分析,基于字形大小构建了3种数据集,并采用PSENet、Pixel Link和PANNet等3种基于分割的深度学习文本检测算法对多种字体的手写藏文古籍文本进行了检测。实验结果显示,3种算法中Pixel Link算法在藏文古籍文献的文本检测中表现最佳,并且Pixel Link算法基于VGG16网络[73]实现了特征提取,展现出在复杂文本检测任务中的显著优势。格桑多吉等[76]提出了一种改进后的轻量级骨干网络Faster-Net作为特征提取网络,引入了注意力机制(Coordinate Attention,CA),充分利用上下文特征信息,实现了浅层与深层信息的有效融合,从而显著提升了对藏文古籍中字体大小不一的文本区域目标检测性能。此外,杨晓龙等[77]提出了一种基于迁移学习的敦煌藏文古籍整页识别方法。该方法有效解决了传统方法在处理敦煌藏文古籍时因数据稀缺和图像复杂性带来的识别率不高的问题。研究通过将预训练的深度学习模型迁移至古籍识别任务中,显著提升了识别精度和模型鲁棒性。该研究不仅在图像预处理和文本区域检测方面取得了良好效果,同时为敦煌藏文古籍的数字化保护和传承提供了有效的技术支持。
KanaarP. Revision of the genus Paratropus Gerstaecker (Coleoptera:Histeri dae)[J]. zoologische verhandelingen, 1997,315(1):1-185.
[3]
KojimaM, NunomiyaC, KawamuraT, et al. Recognition of similar characters by using object oriented design printed Tibetan dictionary[J]. Transaction of Information Processing Society of Japan, 1995,36(11):2611-2621.
[4]
KojimaM, KawazoeY, KimuraM. Automatic Recognition of Wooden Blocked Tibetan Image Character by Using Object Orient ⁃ ed Design[J]. Ipsj Sig Notes,1998,98:39-44.
[5]
RowinskiZ, KeutzerK. Namsel: An optical character recognition system for Tibetan text[J]. Himalayan Lin-guistics,2016:15(1):1544-7502.
[6]
HanY, WangW, WangY, et al. Research on the method of Tibetan recognition based on component locationin-formation[C].Pattern Recognition and Computer Vision: First Chinese Conference, PRCV 2018, Guangzhou,China, Proceedings, Part III 1. Springer International Publishing, 2018:63-73.
[7]
BrodtK, RinchinovO, BazarovA, et al. Deep learning for the development of an OCR for old Tibetan books[C]//Bioinformatics of Genome Regulation and Structure/Systems Biology (BGRS/SB-2022). 2022:1086-1086.
MaS, JinY, ZheJ, et al. A method of printing Tibetan character recognition[C]. Intelligent Control and Automation, Proceedings of the 4th World Congress on. IEEE, 2002.
ZuoJ H, JiaZ H, YangJ, et al. Video image moving target detection based on improved background sub- traction[J].Computer Engineering and Design,2020,401(5):175-180.
MaL, WuJ. A component-based on-line handwritten Tibetan character recognition method using conditional random field[C]. 2012 International Conference on Frontiers in Handwriting Recognition. IEEE, 2012: 704-709.
[40]
MaL, WuJ. Semi-automatic Tibetan component annotation from online handwritten Tibetan character database by optimizing segmentation hypotheses[C]. 12th International Conference on Document Analysis and Recognition. IEEE, 2013:1340-1344.
[41]
MaL, WuJ. A Tibetan component representation learning method for online handwritten Tibetan character recognition[C]. 2014 14th International Conference on Frontiers in Handwriting Recognition. IEEE, 2014:317-322.
ShiB, BaiX, YaoC. An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition[J]. IEEE transactions on pattern analysis and machine intelligence, 2016,39(11):2298-2304.
LiaoM, WanZ, YaoC, et al. Real-time scene text detection with differentiable binarization[C]. Proceedings of the AAAI conference on artificial intelligence, 2020,34(7):11474-11481.
[49]
HowardA, SandlerM, ChuG, et al. Searching for mobilenetv3[C]. Proceedings of the IEEE/CVF international conference on computer vision, 2019:1314-132.
[50]
李金成.藏汉双语自然场景文字检测与识别系统[D].兰州:西北民族大学,2021.
[51]
HeK, ZhangX, RenS, et al. Deep residual learning for image recognition[C]. Proceedings of the IEEE conference on computer vision and pattern recognition, 2016:770-778.
LiaoG, ZhuZ, BaiY, et al.PSENET-based efficient scene text detection[J].IEEE Access,2020,8:52641-52651.
[56]
HeK, ZhangX, RenS, et al.Identity mappings in deep residual networks[C].Computer Vision-ECCV 2016:14th Europe an Conference,Amsterdam,Proceedings,Part IV 14.Springer International Publishing, 2016:630-645.
WangY, WangW, LiZ, et al. Research on Text Line Segmentation of Historical Tibetan Documents Based on the Connected Component Analysis[C]. First Chinese Conference, PRCV 2018, Guangzhou, China, 2018.
[66]
LiZ, WangW, ChenY, et al. A novel method of text line segmentation for historical document image of the uchen Tibetan[J]. Journal of Visual Communication and Image Representation, 2019,61(6):23-32.
[67]
ZhouF, WangW, LinQ. A novel text line segmentation method based on contour curve tracking for Tibetan historical documents[J]. International Journal of Pattern Recognition and Artificial Intelligence, 2018,32(10):1854025.