To promote the learning of entity context and semantic relevance information in text matching, a deep text semantic matching model incorporating entity context features is proposed. This model calculates the comprehensive matching score of text by learning deep multi-view semantic interaction information and entity context feature matching matrix. Bidirectional long short-term memory network (Bi-LSTM) and co-attention mechanism are used to obtain the local semantic features of the text and perform interactive multi-view vector matching. Meanwhile, context features are calculated for the extracted entities in the text, and entity context semantic matching is carried out through entity matching matrix and convolutional neural network. On SNLI, MultiNLI and Quora Question Pairs datasets, the experimental results show that the proposed model can effectively improve the accuracy of text matching compared with the classical deep text matching model.
本文所提出的深度文本语义匹配模型如图2所示,主要包括两个模块:深度多视图语义匹配模块(左)、基于实体上下文特征的语义匹配模块(右)。两个模块同时将语句X:“France is the champion of the 2018 FIFA World Cup”和语句Y:“France won the FIFA World Cup in 2018”作为输入,分别计算文本语义匹配得分和,然后加权求和得到模型最终的综合匹配得分。
在进行实体上下文特征提取之前,要先找到文本里面所包含的命名实体。实体是语句中的一个文本片段,通常由一个或者多个连续的单词构成。如图2中的语句Y“France won the FIFA World Cup in 2018”包含“France”和“the FIFA World Cup”两个实体。直接使用已有自然语言处理基础工具来识别命名实体,如NLTK、StanfordNLP。针对语句X和Y分别抽取得到实体集合和。
假设待匹配的两个语句X和Y的长度分别为m和n,词向量的维度为d,Bi-LSTM隐含层节点数为h,MLP(多层感知机)的隐含层节点数为p,K-Max池化层取k个最大值,为Tensor Layer张量切片数量(slices of tensor),假设从X和Y中分别抽取到了和个实体,那么算法各部分的时间复杂度和空间复杂度如表1所示。由表1可见,本文提出模型的复杂度最大的是Bi-LSTM和向量交互匹配部分,3种相似性交互计算方法中,Tensor Layer的复杂度最高(与张量维度c成正比),Cosine的复杂度最低。
3 实 验
为了验证本文提出的融合实体上下文特征的深度文本语义匹配模型的有效性,分别在SNLI[13]和MultiNLI[14]数据集上测试自然语言推理(natural language inference)任务,在Quora Question Pairs[12]数据集上测试语义鉴别(paraphrase identification)任务。
初始化词向量为预训练的100维GloVe词向量,未登录词(OOV,out of vocabulary)词向量随机初始化,其他可训练变量通过均匀分布随机初始化,Bi-LSTM隐含层维度设为300。全连接网络MLP的隐含层维度设为100,(8)式的非线性激活函数选择tanh函数,dropout概率为0.2,词嵌入层dropout概率为0.5,Tensor Layer中张量的c设为3。深度多视图语义匹配模块的词汇上下文词向量进行一维卷积的卷积核宽度为3。基于实体上下文特征的语义匹配模块中对实体上下文词向量进行一维卷积和的卷积核宽度分别为2和3,对实体匹配矩阵进行二维卷积的卷积核宽度为22。
PANGL, LANY Y, XUJ, et al. A survey on deep text matching[J]. Chinese Journal of Computers,2017,40(4):985-1003. DOI: 10.11897/SP.J.1016.2017.00985(Ch ).
[3]
PARIKHA P, TACKSTROMO, DAS D, et al. A decomposable attention model for natural language inference[C]// Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. Stroudsburg: ACL, 2016:2249-2255.
[4]
YINW, SCHUTZEH. Convolutional neural network for paraphrase identification[C]// Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies. Stroudsburg,PA: NAACL, 2015:901-911.DOI: 10.3115/v1/n15-1091 .
[5]
LIH, XUJ. Semantic matching in search[J]. Foundations and Trends in Information Retrieval, 2014, 7(5): 343-469. DOI: 10.1561/9781601988058 .
YUK, CHENL, CHENB, et al. Cognitive technologies in task-oriented dialogue systems: Concepts, advances and future[J]. Chinese Journal of Computers, 2015, 38(12):2333-2348. DOI: 10.11897/SP.J.1016.2015.02333(Ch ).
[8]
LIUP, QIUX, CHENJ, et al. Deep fusion lstms for text semantic matching[C]// Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics. Stroudsburg: ACL, 2016: 1034-1043.DOI: 10.18653/v1/p16-1098 .
[9]
MIKOLOVT, SUTSKEVERI, CHENK, et al. Distributed representations of words and phrases and their compositionality[C]// Proceedings of the 26th International Conference on Neural Information Processing Systems. Cambridge: MIT Press, 2013: 3111-3119.DOI: 10.5555/2999792.2999959 .
[10]
PENNINGTONJ, SOCHERR, MANNINGC. GloVe: Global vectors for word representation [C]// Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing. Stroudsburg: ACL, 2014:1532-1543.DOI: 10.3115/v1/D14-1162 .
[11]
PETERSM E, NEUMANNM, IYYERM, et al. Deep contextualized word representations[C]// Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics. Stroudsburg: NAACL, 2018:2227-2237.
[12]
DEVLINJ, CHANGM W, LEE K, et al. BERT: pre-training of deep bidirectional transformer for language understanding[C]// Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Stroudsburg: NAACL, 2019: 4171-4186.
[13]
YANGZ L, DAIZ H, YANGY M, et al. XLNet: generalized autoregressive pretraining for language understanding[C]// Proceedings of the 32th International Conference on Neural Information Processing Systems. Cambridge: MIT Press, 2019: 5754-5764.
[14]
AHMADA. Quora question answering dataset[C]// International Conference on Text, Speech, and Dialogue. London: Springer, 2017:66-73.DOI: 10.1007/978-3-319-64206-2_8 .
[15]
BOWMANS R, ANGELIG, POTTSC, et al. A large annotated corpus for learning natural language inference [C]// Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Stroudsburg: ACL, 2015: 632-642.
[16]
WILLIAMSA, NANGIAN, BOWMANS R, et al. A broad-coverage challenge corpus for sentence understanding through inference [C]// Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics. Stroudsburg: NAACL, 2018:1112-1122.
[17]
YANGY, YIH W T, MEEKC. WikiQA: A challenge dataset for open-domain question answering[C]// Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Stroudsburg: ACL, 2015: 2013-2018.
[18]
LANW W, QIUS Y, HEH, et al. A continuously growing dataset of sentential paraphrases[C]// Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. Stroudsburg: ACL, 2017:1224-1234.
[19]
李宏广. 基于深度神经网络的文本匹配算法研究[D]. 合肥: 中国科学技术大学,2019.
[20]
LIH G. Text Matching Based on Deep Neural Network[D]. Hefei: China University of Science and Technology, 2019(Ch).
[21]
HUANGP S, HEX D, GAOJ F, et al. Learning deep structured semantic models for web search using clickthrough data [C]// Proceedings of the 22nd ACM International Conference on Information and Knowledge Management. New York: ACM Press, 2013:2333-2338.DOI: 10.1145/2505515.2505665 .
[22]
SHENY L, HEX D, GAOJ F, et al. A latent semantic model with convolutional-pooling structure for information retrieval[C]// Proceedings the 23rd ACM International Conference on Information and Knowledge Management. New York: ACM Press, 2014:101-110.
[23]
PALANGIH, DENGL, SHENY, et al. Deep sentence embedding using long short-term memory networks: Analysis and application to information retrieval [J]. IEEE Transactions on Audio, Speech, and Language Processing. Piscataway:IEEE,2016, 24(4):694-707. DOI: 10.1109/TASLP.2016.2520371 .
[24]
WANS X, LanY Y, GUOJ F, et al. A deep architecture for semantic matching with multiple positional sentence representations [C]// Proceedings of the 30th AAAI Conference on Artificial Intelligence. Menlo Park: AAAI Press, 2016: 2835-2841.DOI: 10.1007/s10994-013-5363-6 .
[25]
SOCHERR, HUANGE H, PENNINJ, et al. Dynamic pooling and unfolding recursive autoencoders for paraphrase detection[C]// The 25th Annual Conference on Neural Information Processing Systems Neural Information Processing Systems. Cambridge: MIT Press, 2011: 801-809.
[26]
YINW P, SCHUTZEH. MultiGranCNN: An architecture for general matching of text chunks on multiple levels of granularity[C]// Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing. Stroudsburg: ACL, 2015: 63-73.
[27]
LUZ, LIH. A deep architecture for matching short texts[C]// Proceedings of the Advances in Neural Information Processing Systems. Stroudsburg: ACL, 2013: 1367-1375.
[28]
HUB, LUZ, LIH, et al. Convolutional neural network architectures for matching natural language sentences[C]// Proceedings of the Advances in Neural Information Processing Systems. Stroudsburg: ACL, 2014: 2042-2050.
[29]
PANGL, LANY Y, GUOJ F, et al. Text matching as image recognition [C]// Proceedings of the 30th AAAI Conference on Artificial Intelligence. Menlo Park: AAAI Press, 2016: 2793-2799.DOI: 10.5555/3016100.3016292 .
[30]
WANS, LANY, GUOJ, et al. Match-SRNN: Modeling the recursive matching structure with spatial RNN [C]// Proceedings of the 25th International Joint Conference on Artificial Intelligence. San Francisco: Morgan Kaufmann, 2016: 1022-1029.
[31]
KIMS, KANGI, KWAKN, et al. Semantic sentence matching with densely-connected recurrent and co-attentive information[C]// Proceedings of the 33th AAAI Conference on Artificial Intelligence. Menlo Park: AAAI Press, 2019, 33(1):2793-2799. DOI: 10.1609/aaai.v33i01.33016586 .
[32]
GONGY C, LUOH, ZHANGJ, et al. Natural language inference over interaction space[C]// Proceedings of the 6th International Conference on Learning Representations (ICLR). New York: ICLR, 2018:1-15.
[33]
LANW W, XUW . Neural network models for paraphrase identification, semantic textual similarity, natural language inference, and question answering [C]// The 2018 International Conference on Computational Linguistics. New York: ACM Press,2018: 3890-3902 .
[34]
KRIZHEVSKYA, SUTSKEVERI, HINTONG E. Imagenet classification with deep convolutional neural networks [C]// Proceedings of the 26th Annual Conference on Neural Information Processing Systems. Stroudsburg: ACL, 2012: 1097-1105.
[35]
SIMONYANK, ZISSERMANA. Very deep convolutional networks for large-scale image recognition [C]// The 2014 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE,2014:1-14.
[36]
SZEGEDYC, LIUW, JIAY Q, et al. Going deeper with convolutions[C]// The 2015 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway:IEEE Press, 2015: 1-9.
[37]
HUANGG, LIUZ, WEINBERGERK Q, et al. Densely connected convolutional networks [C]// The 2017 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway:IEEE Press, 2017:2261-2269. DOI:10.1109/CVPR.2017.243 .
[38]
HEK, ZHANGX Y, RENS Q, et al. Deep residual learning for image recognition [C]// The 2016 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press,2016: 770-778.
[39]
QIUX, HUANGX. Convolutional neural tensor network architecture for community-based question answering [C]// Proceedings of the 24th International Joint Conference on Artificial Intelligence (IJCAI). San Francisco: Margan Kaufmann, 2015: 1305-1311.
[40]
DUCHIJ, HAZANE, SINGERY. Adaptive subgradient methods for online learning and stochastic optimization [J]. The Journal of Machine Learning Research,2011, 12: 2121-2159.
[41]
CHENQ, ZHUX D, LINGZ H, et al. Enhanced LSTM for natural language inference [C]// Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics. Stroudsburg: ACL. 2017: 1657-1668.DOI: 10.18653/v1/P17-1152 .
[42]
WANGZ G, HAMZAW, FLORIANR. Bilateral multi-perspective matching for natural language sentences [C]// Proceedings of the 26th International Joint Conference on Artificial Intelligence. San Francisco: Morgan Kaufmann, 2017: 4144-4150.DOI: 10.24963/ijcai.2017/579 .
[43]
TAY Y, TUANL A, HUIS C . Compare, compress, and propagate: Enhancing neural architectures with alignment factorization for natural language inference[C]// Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Stroudsburg: ACL, 2018: 1565-1575.DOI: 10.18653/v1/D18-1185 .
[44]
RADFORDA, NARASIMHANK, SALIMANST, et al. Improving Language Understanding by Generative Pre-Training[EB/OL].[2018-06-11].