基于大模型的藏文词性标注方法研究
完么措null , 华却才让 , 环科尤null , 白颖 , 仁青东知
高原科学研究 ›› 2025, Vol. 9 ›› Issue (4) : 129 -136.
基于大模型的藏文词性标注方法研究
Study on Tibetan Part-of-Speech Tagging Methods Based on Large Language Models
针对藏文大语言模型在自然语言处理中的词性标注任务适应性问题,采用标签监督适应方法。以Tibetan-LLaMA2为基础模型,通过深入挖掘其强大的生成能力和丰富的语义表示特性,将最后一层提取的潜在表示映射到词性标签空间,实现精确标注。实验结果显示,无掩码标签监督方法在藏文词性标注数据集中的F1分数达97.1%,并快速适应不同粒度输入(如音节和词),彰显了大模型在复杂语言结构处理上的强大灵活性。
To address the issue of adaptability of the Tibetan large language model to the part-of-speech tagging task in natural language processing, a label-supervised adaptation method was employed in this paper. Based on Tibetan-LLaMA2 as the basic model, by delving into its powerful generative ability and rich semantic representation characteristics, the latent representations extracted at the last layer are mapped to the part-of-speech label space to achieve accurate annotation. Experimental results show that the F1 score of the maskless label-free supervision method in the Tibetan part-of-speech annotation dataset reaches 97.1%, demonstrating its ability to quickly adapt to different granularity inputs (such as syllables and words), and highlighting the strong flexibility of the large language model in processing complex language structures.
| [1] |
刘洋,许乾坤,刘畅, |
| [2] |
|
| [3] |
|
| [4] |
张航,文斌.基于HMM+CRF词性标注的实体抽取方法[J].计算机与数字工程,2023,51(12):2929-2933. |
| [5] |
李亚超,江静,加羊吉, |
| [6] |
格桑加措.基于HMM模型的藏语词性标注研究[J].信息通信,2020(5):46-47. |
| [7] |
康才畯.藏语分词与词性标注研究[D].上海:上海师范大学,2014. |
| [8] |
王康.基于神经网络的藏语分词与词性标注研究[D].兰州:兰州大学,2020. |
| [9] |
索朗次仁.基于预训练语言模型的藏文分词与词性标注研究[D].拉萨:西藏大学,2023. |
| [10] |
杨毛加,柔特,才智杰, |
| [11] |
|
| [12] |
|
| [13] |
|
| [14] |
袁里驰.基于BERT-BiLSTM-CRF的中文分词和词性标注联合方法[J].小型微型计算机系统,2023,44(9):1906-1911. |
| [15] |
国家市场监督管理总局. 全国信息技术标准化技术委员会信息处理用藏语词类标记集( GB/T 36337-2018)[S].北京:中国标准出版社,2018. |
| [16] |
常博林,袁义国,李斌, |
国家自然科学基金项目(62166034)
藏语智能全国重点实验室建设项目(2025-ZJ-J08)
/
| 〈 |
|
〉 |