1.School of Computer Science and Technology,Wuhan University of Science and Technology,Wuhan 430065,Hubei,China
2.Big Data Science and Engineering Research Institute,Wuhan University of Science and Technology,Wuhan 430065,Hubei,China
3.Hubei Province Key Laboratory of Intelligent Information Processing and Real-Time Industrial System,Wuhan 430065,Hubei,China
4.The Key Laboratory of Rich-Media Knowledge Organization and Service of Digital Publishing,National Press and Publication Administration,Beijing 100038,China
5.Center for Evidence-Based and Translational Medicine,Zhongnan Hospital of Wuhan University,Wuhan 430071,Hubei,China
In recent years, the publication speed of medical guidelines has been rapid, but their structured organization and retrieval in the evidence-based process face many challenges, which affect the rapid acquisition and application of knowledge. This study proposes a guideline knowledge extraction model based on evidence-based medicine to address the issue. This model aims to efficiently extract guideline evidence and identify the risk of bias in the source literature of guideline evidence. The model combines manually annotated small-sample datasets to construct prompt instructions for training large language models to improve their performance in guideline knowledge extraction tasks. A randomized controlled trial bias risk assessment model was constructed that integrated external knowledge, and enhanced the model's bias risk identification ability by utilizing diverse external knowledge information based on the BERT model. After fine-tuning with prompts, the performance of the large language model in extracting key information from guidelines was significantly improved, with BLEU-4 improving by over 43 percentage points. The evidence bias risk assessment model that integrates external knowledge also showed superior performance in identifying bias risks, with a precision of 0.898 and an F1 value of 0.892. With a limited dataset size, fine-tuning a large language model with hint engineering can effectively extract the required medical guideline information. The incorporation of various external knowledge sources substantially improved the effectiveness of the bias risk identification model. This suggests that models designed to extract knowledge from medical guidelines can contribute to enhancing the quality of clinical research.
因此,系统评价需要对纳入的RCT进行质量分析,排除低质量的RCT或对不同偏倚风险的RCT结果进行对比分析。Cochrane系统评价手册[4]中提到,RCT的偏倚风险类别主要从以下7个方面进行评估:随机序列生成(Random Sequence Generation,RSG)、分配隐藏(Allocation Concealment,AC)、对试验受试者及试验人员实施盲法(Blinding Of Participants and personnel,BOP)、对结局评估员实施盲法(Blinding of Outcome Assessment,BOA)、实验结果数据不完整(Incomplete Outcome Data,IOD)、选择性报告(Selective Reporting,SR)和其他偏倚(Other sources of bias,O)。评估的结果用“偏倚风险较低”“风险未知”或“偏见风险较高”表示。另外,有研究表明,大部分RCT的偏倚风险评价需要花费系统评估员10到60分钟的时间才能完成,并且偏倚风险判断结果可能还不太准确。如果遗漏了文献的一些关键信息,不同的系统评估员针对同一篇RCT甚至可能得出不一样的结论[5]。
当前,有很多用于偏倚风险评估的工具,比如Cochrane ROB tool、PEDro量表、Delphi清单、Jadad量表等等,但是这些工具极其依赖人工操作,既耗时又耗力。而且随着互联网和科学技术的发展,近些年来生物医学文献的发表数量呈现爆发式增长,对其进行偏倚风险评估变得更加困难。因此,如果采用人工智能的方法实现自动化偏倚风险评估,将是一项十分有意义的研究,不仅能在一定程度上减轻系统评估员的工作量,同时也可以促进快速医疗决策。
大语言模型(Large Language Model,LLM)是一种集成了巨量数据的语言模型,其数据参数规模高达亿级别[7]。它的主要目的是解析人类的问题并提供精确的回应。通过其强大的数据处理能力,LLM能够预测文本序列中的后续词汇,或生成与给定文本紧密相关的内容。这使得LLM在自然语言处理领域具有广泛的应用价值,包括但不限于文本分类、情感分析、智能问答以及机器翻译等。简而言之,LLM能够理解人类的询问并提供精准的回应,为我们提供便利。
本文使用的数据主要来自Cochrane系统评价数据库(Cochrane Database of Systematic Reviews,CDSR)[19]。CDSR几乎涵盖了临床医学的各个专业,截至2023年底,已经包含了9 100余篇系统评价、2 300余篇研究方案计划书和部分社论等。
系统评价包含的结构化偏倚风险评价数据主要分为三个部分:Bais、Authors' judgement和Support for judgement。Bais也就是偏倚风险的类别。Authors' judgement是作者给出的对该类偏倚风险的判断结果,低偏倚风险Low Risk、高偏倚风险High Risk或者偏倚风险未知Unclear Risk。Support for judgement是作者对此判断结果给出的依据。
CHENY L, SUNY J, LUOX F,et al. The core methods and key models in evidence-based medicine[J]. Medical Journal of Peking Union Medical College Hospital, 2023,14(1):1-8. DOI: 10.12290/xhyxzz.2022-0686 .
LUOH, LIUJ P. Quality and validity of randomized controlled trials in China from the perspective of systematic reviews[J]. Journal of Chinese Integrative Medicine, 2011, 9(7): 697-701 (Ch). DOI: 10.3736/jcim20110701 .
[5]
胡雁. 循证护理学[M]. 北京:人民卫生出版社,2012.
[6]
HUY. Evidence based nursing[M].Beijing:People’s Medical Publishing House,2012 (Ch).
[7]
SAVOVIĆJ, WEEKSL, STERNEJ A C, et al. Evaluation of the Cochrane Collaboration’s tool for assessing the risk of bias in randomized trials: Focus groups, online survey, proposed recommendations and their implementation[J]. Systematic Reviews, 2014, 3: 37. DOI: 10.1186/2046-4053-3-37 .
[8]
LENSENS, FARQUHARC, JORDONV. Risk of bias: Are judgements consistent between reviews[DB/OL].[2024-01-02].
[9]
MARSHALLI J, KUIPERJ, WALLACEB C. Automating risk of bias assessment for clinical trials[C]//Proceedings of the 5th ACM Conference on Bioinformatics, Computational Biology, and Health Informatics. Newport Beach California. New York: ACM, 2014: 88-95. DOI: 10.1145/2649387.2649406 .
ZHUW, LIY G, LIZ, et al. Research on large language model of intensive care medicine based on literature and knowledge learning[J]. China Digital Medicine, 2024, 19(3): 36-41 (Ch). DOI: 10.3969/j.issn.1673-7571.2024.03.007 .
[12]
LIUX, JIK X, FUY C, et al. P-Tuning v2: Prompt tuning can be comparable to fine-tuning across scales and tasks[DB/OL].[2023-01-02]. DOI: 10.18653/v1/2022.acl-short.8 .
[13]
LOUKASL, STOGIANNIDISI, MALAKASIOTISP, et al. Breaking the bank with ChatGPT: Few-shot text classification for finance[DB/OL].[2023-01-02]. DOI: 10.1145/3604237.3626891 .
LIX Y, QIANL, ZHANGZ X. A scoring method for semantic evaluation metrics of scientific papers based on prompt tuning of large language models[J]. Data Analysis and Knowledge Discovery, 2024, 8(Z1): 200-212 (Ch).
[16]
ZHANGY, JINR, ZHOUZ H. Understanding bag-of- words model:A statistical framework[J].International Journal of Machine Learning and Cybernetics,2010,1(1):43-52. DOI:10.1007/s13042-010-0001-0 .
[17]
JIANGZ Y, GAOB, HEY L,et al. Text classification using novel term weighting scheme-based improved TF-IDF for Internet media reports[J].Mathematical Problems in Engineering,2021: Article ID 6619088. DOI: 10.1155/2021/6619088 .
[18]
MILLARDL A C, FLACHP A, HIGGINSJ P T. Machine learning to assist risk-of-bias assessments in systematic reviews[J]. International Journal of Epidemiology, 2016, 45(1): 266-277. DOI: 10.1093/ije/dyv306 .
XIAY. Research on relevant algorithms for automated risk of bias assessment in systematic reviews[D].Chengdu: University of Electronic Science and Technology of China, 2020(Ch).
[21]
DEVLINJ, CHANGM W, LEEK, et al. BERT: Pre-training of deep bidirectional transformers for language understanding[EB/OL]. 2018: arXiv: 1810.04805. DOI: 10.48550/arXiv.1810.04805 .
[22]
LIANGD K, XUW, BAIX. An end-to-end transformer model for crowd localization[EB/OL]. 2022: arXiv: 2202.13065. DOI: 10.1007/978-3-031-19769-7_3 .
SUNH, SHIJ P. Expansion of PICO model in evidence-based medicine and its application in qualitative research[J]. Chinese Journal of Evidence-Based Medicine, 2014, 14(5): 505-508. DOI: 10.7507/1672-2531.20140087 (Ch ).
BAIY M, DUJ. Computable clinical evidence synthesis: A literature review[J]. Journal of Capital Medical University, 2022, 43(4): 576-583. DOI: 10.3969/j.issn.1006-7795.2022.04.011 (Ch ).
[29]
WANGQ, LIAO JY, LAPATAM, et al. Risk of bias assessment in preclinical literature using natural language processing[J]. Res Synth Methods, 2022, 13(3): 368-380. DOI: 10.1002/jrsm.1533 .