PDF (596K)
摘要
为解决传统的mT5模型在提取文本关键信息方面存在局限性,无法准确捕获长文本中的所有关键内容,导致生成的摘要信息缺乏全面性和准确性等问题,提出一种基于mT5的融合注意力机制与LSTM的生成式文本摘要方法(mT5-ATT-LSTM).该方法首先利用mT5的编码器获取输入文本的初始词向量表示;然后引入注意力机制增强词向量表示,使模型能够关注输入文本中的关键信息;接着通过引入LSTM来捕捉长程依赖关系;最后通过mT5的解码器将信息转换为目标文本生成摘要.在公开的中文新闻短文本数据集LCSTS上对本方法进行实验,本方法的ROUGE-1、ROUGE-2、ROUGE-L分别为37.71、24.61、34.51,相比传统mT5模型ROUGE值分别提升了2.1%、3.3%、2.1%,通过对比和消融实验验证了该方法的有效性.
Abstract
To address the limitations of traditional mT5 models in extracting key textual information, particularly their inability to accurately capture all critical content within lengthy texts, thereby resulting in generated summaries lacking comprehensiveness and accuracy, we propose a generative text summarization approach integrating attention mechanism and Long Short Term Memory (LSTM), denoted as mT5-ATT-LSTM. This approach first utilizes mT5 encoder to obtain initial word vector representations of input text. Subsequently, an attention mechanism is introduced to enhance word vector representations, enabling the model to focus on different aspects of the content. Then, LSTM is employed to capture long-range dependency relationships. Finally, mT5 decoder is utilized to transform the information into target text for summary generation. Experimental evaluations conducted on the publicly available Chinese news short text dataset LCSTS demonstrate the effectiveness of the proposed method. The ROUGE-1, ROUGE-2, and ROUGE-L scores achieved by this method are 37.71, 24.61, and 34.51, respectively. Compared to the traditional mT5 model, our approach yields improvements of 2.1%, 3.3%, and 2.1% in ROUGE scores, validating the efficacy of the proposed method through comparative and ablation experiments.
关键词
Key words
[Author(id=1305171447027040558, tenantId=1045748351789510663, journalId=1155139928303341810, articleId=1305171446062350582, orderNo=0, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=null, emailSecond=null, emailThird=null, correspondingAuthor=0, authorType=1, ext={EN=AuthorExt(id=1305171447085760820, tenantId=1045748351789510663, journalId=1155139928303341810, articleId=1305171446062350582, authorId=1305171447027040558, language=EN, stringName=Feng LIU, firstName=Feng, middleName=null, lastName=LIU, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=School of Computer Science, Hubei Univ. of Tech. , Wuhan 430068, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1305171447127703863, tenantId=1045748351789510663, journalId=1155139928303341810, articleId=1305171446062350582, authorId=1305171447027040558, language=CN, stringName=刘冯, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=湖北工业大学 计算机学院 , 湖北武汉 430068, bio={"content":"刘冯(1999-),男,湖北武汉人,湖北工业大学硕士研究生,研究方向为自然语言处理文本摘要算法.
"}, bioImg=null, bioContent=刘冯(1999-),男,湖北武汉人,湖北工业大学硕士研究生,研究方向为自然语言处理文本摘要算法.
, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1305171446955737385, tenantId=1045748351789510663, journalId=1155139928303341810, articleId=1305171446062350582, xref=null, ext=[AuthorCompanyExt(id=1305171446972514602, tenantId=1045748351789510663, journalId=1155139928303341810, articleId=1305171446062350582, companyId=1305171446955737385, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=School of Computer Science, Hubei Univ. of Tech. , Wuhan 430068, China), AuthorCompanyExt(id=1305171446985097515, tenantId=1045748351789510663, journalId=1155139928303341810, articleId=1305171446062350582, companyId=1305171446955737385, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=湖北工业大学 计算机学院 , 湖北武汉 430068)])]), Author(id=1305171447169646907, tenantId=1045748351789510663, journalId=1155139928303341810, articleId=1305171446062350582, orderNo=1, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=null, emailSecond=null, emailThird=null, correspondingAuthor=1, authorType=1, ext={EN=AuthorExt(id=1305171447224172863, tenantId=1045748351789510663, journalId=1155139928303341810, articleId=1305171446062350582, authorId=1305171447169646907, language=EN, stringName=Caiquan XIONG, firstName=Caiquan, middleName=null, lastName=XIONG, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=School of Computer Science, Hubei Univ. of Tech. , Wuhan 430068, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1305171447270310211, tenantId=1045748351789510663, journalId=1155139928303341810, articleId=1305171446062350582, authorId=1305171447169646907, language=CN, stringName=熊才权, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=湖北工业大学 计算机学院 , 湖北武汉 430068, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1305171446955737385, tenantId=1045748351789510663, journalId=1155139928303341810, articleId=1305171446062350582, xref=null, ext=[AuthorCompanyExt(id=1305171446972514602, tenantId=1045748351789510663, journalId=1155139928303341810, articleId=1305171446062350582, companyId=1305171446955737385, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=School of Computer Science, Hubei Univ. of Tech. , Wuhan 430068, China), AuthorCompanyExt(id=1305171446985097515, tenantId=1045748351789510663, journalId=1155139928303341810, articleId=1305171446062350582, companyId=1305171446955737385, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=湖北工业大学 计算机学院 , 湖北武汉 430068)])])]
刘冯,熊才权.
基于mT5的融合注意力机制与LSTM的生成式文本摘要方法研究[J].
湖北工业大学学报, 2026, 41(4): 70-78 DOI:
| [1] |
Conroy J M, O'leary D P. Text summarization via hidden markov models[C]// Proceedings of the 24th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 2001: 406-407.
|
| [2] |
Shen D, Sun J T, Li H, et al. Document summarization using conditional random fields[C]// IJCAI. 2007, 7: 2862-2867.
|
| [3] |
Mihalcea R, Tarau P. Textrank: Bringing order into text[C]// Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing. 2004: 404-411.
|
| [4] |
Hennig L. Topic-based multi-document summarization with probabilistic latent semantic analysis[C]// Proceedings of the International Conference RANLP-2009. 2009: 144-149.
|
| [5] |
Sutskever I, Vinyals O, Le Q V. Sequence to sequence learning with neural networks[J]. Advances in Neural Information Processing Systems, 2014, 27: 3104-3112.
|
| [6] |
Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[J]. Advances in Neural Information Processing Systems, 2017, 30: 5998-6008.
|
| [7] |
Brown T, Mann B, Ryder N, et al. Language models are few-shot learners[J]. Advances in Neural Information Processing Systems, 2020, 33: 1877-1901.
|
| [8] |
Devlin J, Chang M W, Lee K, et al. Bert: Pre-training of deep bidirectional transformers for language understanding[J]. arXiv preprint arXiv:1810.04805, 2018.
|
| [9] |
Raffel C, Shazeer N, Roberts A, et al. Exploring the limits of transfer learning with a unified text-to-text transformer[J]. The Journal of Machine Learning Research, 2020, 21(1): 5485-5551.
|
| [10] |
Xue L, Constant N, Roberts A, et al. mT5: A massively multilingual pre-trained text-to-text transformer[J]. arXiv preprint arXiv:2010.11934, 2020.
|
| [11] |
Abadi V N M, Ghasemian F. Enhancing Persian text summarization through a three-phase fine-tuning and reinforcement learning approach with the mT5 transformer model[J]. Scientific Reports, 2025, 15(01): 80.
|
| [12] |
Hasan T, Bhattacharjee A, Islam M S, et al. XL-sum: Large-scale multilingual abstractive summarization for 44 languages[J]. arXiv preprint arXiv:2106.13822, 2021.
|
| [13] |
Calizzano R, Ostendorff M, Ruan Q, et al. Generating Extended and Multilingual Summaries with Pre-trained Transformers[C]// Proceedings of the Thirteenth Language Resources and Evaluation Conference. 2022: 1640-1650.
|
| [14] |
Yang Z G. Neural text summarization for Hungarian[J]. Acta Linguistica Academica, 2022, 69(04): 474-500.
|
| [15] |
Mukherjee A. Developing Bengali Text Summarization with Transformer Base model[D]. Dublin: National College of Ireland, 2022.
|
| [16] |
Gryaznov A, Rybka R, Moloshnikov I, et al. Influence of the duration of training a deep neural network model on the quality of text summarization task[C]// AIP Conference Proceedings. AIP Publishing LLC, 2023, 2849(01): 400006.
|
| [17] |
Chi Z, Dong L, Ma S, et al. mT6: Multilingual pre-trained text-to-text transformer with translation pairs[J]. arXiv preprint arXiv:2104.08692, 2021.
|
| [18] |
Galeshchuk S. Abstractive Summarization for the Ukrainian Language: Multi-Task Learning with Hromadske.ua News Dataset[C]// Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP). 2023: 49-53.
|
| [19] |
lincy. ROUGE: A package for automatic evaluation of summaries[C]// ACL 2004: Text Summarization Branches Out, 2004: 74-81.
|
| [20] |
Hu B, Chen Q, Zhu F. Lcsts: A large scale Chinese short text summarization dataset[J]. arXiv preprint arXiv:1506.05865, 2015.
|
| [21] |
崔卓, 李红莲, 张乐, 等. . 一种融合义原的中文摘要生成方法[J]. 中文信息学报, 2022, 36(06): 146-154.
|
| [22] |
Li Z, Wu J, Miao J, et al. A topic inference Chinese news headline generation method integrating copy mechanism[J]. Neural Processing Letters, 2023, 55(2): 1337-1353.
|
| [23] |
Devlin J, Chang M W, Lee K, et al. Bert: Pre-training of deep bidirectional transformers for language understanding[J]. arXiv preprint arXiv:1810.04805, 2018.
|
| [24] |
Zhuang L, Wayne L, Ya S, et al. A robustly optimized BERT pre-training approach with post-training[C]// Proceedings of the 20th Chinese national conference on computational linguistics. 2021: 1218-1227.
|
| [25] |
Ma S, Sun X, Lin J, et al. Autoencoder as assistant supervisor: improving text representation for Chinese social media text summarization[C]// Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Melboume: ACL Press, 2018: 725-731.
|
基金资助
湖北省科技计划项目(2021BLB171)
国家自然科学基金(61902116)