The industrial dataset comprises various types of operation documents, maintenance manuals, and equipment drawings, as well as an increasing number of work orders and job records. Existing general-term extraction methods are limited in effectiveness in Chinese contexts, and the scarcity of prior resources also makes traditional supervised learning processes challenging to implement. Therefore, standard models in the industry do not perform ideally in the task of extracting vertical domain terminology in industrial fields. This study introduces a dual-step extraction strategy based on pre-extraction and fine-tuning refinement to address these issues. the approach enhances the model's ability to capture semantic information using the XLNet pre-trained model and incorporating character, glyph, and phonetic features. Moreover, the paper employs an LSTM encoder-decoder model to generate negative samples containing typographical errors to expand the dataset, aiming to improve the model's robustness to noisy text. Applying the method in this article to the field of the automotive industry,experimental results show that our method improves performance in the industrial vertical domain by 17% over the existing traditional methods, demonstrating its effectiveness.
术语是专业领域中为了准确表达专业概念和指称专业对象而创造的专业(科技及其他)语言中的词或词组,对于自然语言处理工作的下游任务产生着重要的影响。自动术语抽取(Automatic Term Extraction,ATE)是自然语言处理领域内的核心课题,旨在从专业领域特定文本中识别和提取专业术语。如工业领域中的专业术语有机械结构、零件名称和工艺流程等,这些术语对于理解和处理该领域的文本至关重要。本研究的目标是通过学习有限规模的专业性文本语料库,抽取涵盖机械结构、零件名称和工艺流程等汽车工业领域的专业术语。
BERT是一种基于自编码的预训练模型[5],使用掩码语言模型(Masked Language Model,MLM)进行训练。这种方法存在预训练和微调之间的不匹配,因为在预训练阶段,模型需要预测被掩码的词,而在微调阶段,模型处理的是未掩码的输入。与之相反,XLNet采用了一种自回归方法,同时保留了BERT的双向上下文建模能力。XLNet使用置换语言建模(PLM)对所有可能的因子分解顺序进行建模,从而解决了预训练和微调阶段的不匹配问题。
本文自建的数据集AMID (Automotive Manufacturing Industry Data)分为两个部分,高质量语料库,示例如图8;低质量语料库,示例如图9。高质量语料库的语言逻辑规整,术语密集且重复出现,抽取难度较小的高质量语料库。主要由工业汽车制造领域的用户手册、技术文档、技术图纸组成。
KAGEURAK, UMINOB. Methods of automatic term recognition: A review[J]. Terminology. International Journal of Theoretical and Applied Issues in Specialized Communication, 1996, 3(2): 259-289. DOI: https://doi.org/10.1075/term.3.2.03kag .
[2]
ASTRAKHANTSEVN. ATR4S: Toolkit with state-of-the-art automatic terms recognition methods in Scala[J]. Language Resources and Evaluation, 2018, 52(3): 853-872. DOI: 10.1007/s10579-017-9409-4 .
ZHANGX, SUNH Y, XIND X, et al. Survey on automatic term extraction research[J]. Journal of Software, 2020, 31(7): 2062-2094. DOI: 10.13328/j.cnki.jos.006040 (Ch ).
[5]
TRANH T H, MARTINCM, CAPORUSSOJ, et al. The recent advances in automatic term extraction: A survey[EB/OL]. 2023: arXiv: 2301.06767.
[6]
DEVLINJ, CHANGM W, LEEK, et al. BERT: Pre-training of deep bidirectional transformers for language understanding[EB/OL]. 2018: arXiv: 1810.04805. DOI: 10.48550/arXiv.1810.04805 .
[7]
CUIY M, CHEW X, LIUT, et al. Pre-training with whole word masking for Chinese BERT[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 29: 3504-3514. DOI:10.1109/TASLP.2021.3124365 .
[8]
SUNY, WANGS H, LIY K, et al. ERNIE: Enhanced representation through knowledge integration[EB/OL]. 2019:arXiv:1904.09223. DOI:10.48550/arXiv.1904.09223 .
[9]
DIAOS, BAIJ, SONGY, et al. ZEN: Pre-training Chinese text encoder enhanced by n-gram representations[EB/OL]. arXiv: 2019:1911.00720. DOI:10.48550/arXiv.1911.00720 .
[10]
FRANTZIK, ANANIADOUS, MIMAH. Automatic recognition of multi-word terms: The C-value/NC-value method[J]. International Journal on Digital Libraries, 2000, 3(2): 115-130. DOI: 10.1007/s007999900023 .
[11]
ZHANGX W, WUP, CAIJ M, et al. A contrastive study of Chinese text segmentation tools in marketing notification texts[J]. Journal of Physics: Conference Series, 2019, 1302(2): 022010. DOI: 10.1088/1742-6596/1302/2/022010 .
[12]
LUOR X, XUJ J, ZHANGY, et al. PKUSEG: A toolkit for multi-domain Chinese word segmentation[EB/OL]. 2019: arXiv: 1906.11455.
[13]
MAJ, GANCHEVK, WEISSD. State-of-the-art Chinese word segmentation with Bi-LSTMs[EB/OL]. 2018: arXiv: 1808.06511. DOI: 10.18653/v1/d18-1529 .
[14]
MIKOLOVT, CHENK, CORRADOG, et al. Efficient estimation of word representations in vector space[EB/OL]. 2013: arXiv: 1301.3781.
[15]
CAMPOSR, MANGARAVITEV, PASQUALIA, et al. YAKE! Keyword extraction from single documents using multiple local features[J]. Information Sciences, 2020, 509: 257-289. DOI: 10.1016/j.ins.2019.09.013 .
[16]
RASMUSA, BERGLUNDM, HONKALAM, et al. Semi-supervised learning with ladder networks[DB/OL].[2023-05-16].
CUIY M, CHEW X, LIUT, et al. Revisiting pre-trained models for Chinese natural language processing[C]//Findings of the Association for Computational Linguistics. Stroudsburg: Association for Computational Linguistics, 2020: 657-668. DOI: 10.18653/v1/2020.findings-emnlp.58 .
[26]
RIGOUTS TERRYNA, HOSTEV, DROUINP, et al. TermEval 2020: Shared task on automatic term extraction using the annotated corpora for term extraction research (acter) dataset[DB/OL]. [2023-04-30].
[27]
ROCHETEAUJ, DAILLEB. TTC TermSuite: A UIMA application for multilingual terminology extraction from comparable corpora[DB/OL].[2023-05-16]. DOI: 10.1007/11562214_62 .
WUJ, CHENGY, HAOH, et al. Automatic extraction of Chinese terminology based on BERT embedding and BiLSTM-CRF model[J]. Journal of the China Society for Scientific and Technical Information, 2020, 39(4):409-418. DOI:10.3772/j.issn.1000-0135.2020.04.007 (Ch ).
[30]
HANX W, XUL Z, QIAOF. CNN-BiLSTM-CRF model for term extraction in Chinese corpus[C]//International Conference on Web Information Systems and Applications. Cham: Springer, 2018: 267-274. DOI:10.1007/978-3-030-02934-0_25 .
[31]
TERRYNA R, HOSTEV, LEFEVERE. HAMLET: Hybrid adaptable machine learning approach to extract terminology[J]. Terminology. International Journal of Theoretical and Applied Issues in Specialized Communication, 2021, 27(2): 254-293. DOI:10.1075/term.20017.rig .
[32]
SUNC, QIUX P, XUY G, et al. How to fine-tune BERT for text classification[C]// Chinese Computational Linguistics. Cham: Springer, 2019: 194-206. DOI:10.1007/978-3-030-32381-3_16 .
[33]
SOUZAF, NOGUEIRAR, LOTUFOR. Portuguese named entity recognition using BERT-CRF[EB/OL]. 2019: arXiv: 1909.10649.