Department of Epidemiology and Health Statistics, Xiangya School of Public Health, Central South University, Changsha 410013, China
Show less
文章历史+
Received
Published
2024-12-05
2025-07-28
Issue Date
2025-10-30
PDF (1369K)
摘要
目的 性早熟危险因素的准确识别有助于临床诊疗,但运用自然语言处理非结构化数据的方法仍有待评价。本研究旨在基于性早熟电子病历中个体危险因素抽取评价提示词工程方法的性能。 方法 根据CRISPE(capacity and role-insight-statement-personality-experiment)提示词框架制订简单提示词和优化提示词,2种提示词分别引导大语言模型GLM-4-9B从653份电子病历记录中提取10种性早熟的危险因素,采用准确率、精确率、召回率和F1值作为信息抽取任务的评价指标。 结果 在简单提示词和优化提示词下,模型总体的准确率、精确率、召回率和F1值分别为84.18%、98.09%、81.99%、89.32%和97.15%、98.31%、98.16%、98.23%。优化提示词在年龄(<9岁和≥9岁)和就诊时间(<2023年和≥2023年)各组间的模型性能差异小于简单提示词。在简单提示词下,模型抽取每种危险因素的准确率的区间范围为60.03%~97.24%;在优化提示词下,准确率的区间范围为92.19%~99.85%。2种提示词在抽取“饮料摄入情况”时的准确率差异最大(60.03% vs 92.19%),在抽取“母亲初潮年龄”时差异最小(97.24% vs 99.23%)。在简单提示词、优化提示词和真实值3种情况下,零食摄入情况、饮料摄入情况、豆浆摄入情况、蜂蜜摄入情况、保健品服用情况、补品服用情况、睡眠质量、开灯睡觉情况的分布特征差异均具有统计学意义(均P<0.001),运动情况(P=0.966)和母亲初潮年龄(P=0.952)的分布特征差异无统计学意义。 结论 优化提示词相比简单提示词更能有效地完成电子病历中个体危险因素的抽取任务,表明提示词工程在提升大语言模型性能方面具有重要作用。
Abstract
Objective Accurate identification of risk factors for precocious puberty is essential for clinical diagnosis and management, yet the performance of natural language processing methods applied to unstructured electronic medical record (EMR) data remains to be fully evaluated. This study aims to assess the performance of a prompt engineering method for extracting individual risk factors of precocious puberty from EMRs. Methods Based on the capacity and role-insight-statement-personality-experiment (CRISPE) prompt framework, both simple and optimized prompts were designed to guide the large language model GLM-4-9B in extracting 10 types of risk factors for precocious puberty from 653 EMRs. Accuracy, precision, recall, and F1-score were used as evaluation metrics for the information extraction task. Results Under simple and optimized prompt conditions, the overall accuracy, precision, recall, and F1-score of the model were 84.18%, 98.09%, 81.99%, and 89.32% versus 97.15%, 98.31%, 98.16%, and 98.23%, respectively. The optimized prompts achieved more stable performance across age (<9 years vs ≥9 years) and visit-time (<2023 vs ≥2023) subgroups compared with simple prompts. The accuracy range for extracting each risk factor was 60.03%-97.24%, while with optimized prompts, the range improved to 92.19%-99.85%. The largest performance improvement occurred for “beverage intake” (60.03% vs 92.19%), and the smallest for “maternal age of menarche” (97.24% vs 99.23%). In comparing distributions among simple prompts, optimized prompts, and ground truth, statistically significant differences were observed for snack intake, beverage intake, soy milk intake, honey intake, supplement use, tonic use, sleep quality, and sleeping with the light on (all P<0.001), while exercise (P=0.966) and maternal menarche age (P=0.952) showed no significant differences. Conclusion Compared with simple prompts, optimized prompts substantially improved the extraction performance of individual risk factors for precocious puberty from EMRs, underscoring the critical role of prompt engineering in enhancing large language model performance.
首先,设计一组简单提示词引导大语言模型明确任务要求并识别个体危险因素,实体关系抽取设定输出数据类型为半结构化的JavaScript对象表示法(JavaScript object notation,JSON)格式。其次,根据CRISPE(capacity and role-insight-statement-personality-experiment)提示词框架优化提示词[10],相较于简单提示词增加了角色和能力、上下文信息及补充信息等内容,旨在提升大语言模型对病历文本中关键信息的敏感度。简单提示词和优化提示词的内容比较见表3。
The Subspecialty Group of Endocrinologic, Hereditary and Metabolic Diseases, the Society of Pediatrics, Chinese Medical Association, the Editorial Board, Chinese Journal of Pediatrics, FUJunfen, et al. Expert consensus on the diagnosis and treatment of central precocious puberty(2022)[J]. Chinese Journal of Pediatrics, 2023, 61(1): 16-22.
[3]
LiuYF, YuTT, LiXQ, et al. Prevalence of precocious puberty among Chinese children: a school population-based study[J]. Endocrine, 2021, 72(2): 573-581.
[4]
DongY, DaiLL, DongY, et al. Analysis of risk factors of precocious puberty in children[J]. BMC Pediatr, 2023, 23(1): 456.
JIXurui, WEIDejian, ZHANGJunzhong, et al. Research progress on information extraction methods of Chinese electronic medical records[J]. Computer Engineering & Science, 2024, 46(2): 325-337.
[7]
NashwanAJ, AbuJaberAA. Harnessing the power of large language models (LLMs) for electronic health records (EHRs) optimization[J/OL]. Cureus, 2023, 15(7): e42634[2024-09-30].
WANGDongqing, LUFei, ZHANGBinghui, et al. Survey on prompt engineering in large language model[J]. Computer Systems and Applications, 2025, 34(1): 1-10.
YANGBo, LIUQin, LIUShudan, et al. Relevant factors of early puberty timing: a systematic review[J]. Chinese Journal of Evidence-Based Medicine, 2018, 18(12): 1337-1351.
[12]
蔡璐. 学龄期女童性早熟的相关危险因素分析[D]. 大连: 大连医科大学, 2022.
[13]
CAILu. Risk analysis of precocious puberty in school-aged girls[D]. Dalian: Dalian Medical University, 2022.
[14]
LiY, GaoD, ChenM, et al. Association between healthy lifestyle pattern and early onset of puberty: based on a longitudinal follow-up study[J]. Br J Nutr, 2022, 128(12): 2320-2329.
[15]
ZhongR, XuY, ZhangC, et al. Leveraging large language model to generate a novel metaheuristic algorithm with CRISPE framework[J]. Clust Comput, 2024, 27(10): 13835-13869.
[16]
ZengA, XuB, WangB, et al. ChatGLM: A family of large language models from GLM-130B to GLM-4 all tools[EB/OL]. (2024-08-09)[2024-09-29].
[17]
BartalesiV, LenziE, De MartinoC. Using large language models to create narrative events[J/OL]. Peer J Comput Sci, 2024, 10: e2242[2024-09-20].
[18]
GuQY, WuYM, FengZW, et al. Dietary pattern and precocious puberty risk in Chinese girls: a case-control study[J]. Nutr J, 2024, 23(1): 14.
[19]
PatelS, RahmaniB, GandhiJ, et al. Revisiting the pineal gland: a review of calcification, masses, precocious puberty, and melatonin functions[J]. Int J Neurosci, 2020, 130(5): 464-475.
[20]
HartsteinLE, Diniz BehnC, WrightKP, et al. Evening light intensity and phase delay of the circadian clock in early childhood[J]. J Biol Rhythms, 2023, 38(1): 77-86.
[21]
CaseyJA, SchwartzBS, StewartWF, et al. Using electronic health records for population health research: a review of methods and applications[J]. Annu Rev Public Health, 2016, 37: 61-81.
GENGGuojun, WANGLuming, TANGMingkun, et al. Chinese expert consensus on quality control and management of electronic medical records for thoracic surgery(2024 version)[J/OL]. Chinese Journal of Clinical Thoracic and Cardiovascular Surgery, 2024: 1-10[2024-11-07].
[24]
WangL, ChenX, DengXW, et al. Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs[J]. NPJ Digit Med, 2024, 7(1): 41.
[25]
MeskóB. Prompt engineering as an important emerging skill for medical professionals: tutorial[J/OL]. J Med Internet Res, 2023, 25: e50638[2024-11-11].
[26]
FinkMA, BischoffA, FinkCA, et al. Potential of ChatGPT and GPT-4 for data mining of free-text CT reports on lung cancer[J/OL]. Radiology, 2023, 308(3): e231362[2024-11-13].
[27]
ChenYL, HuangXQ, TianL. Meta-analysis of machine learning models for the diagnosis of central precocious puberty based on clinical, hormonal (laboratory) and imaging data[J]. Front Endocrinol (Lausanne), 2024, 15: 1353023.
ZHOUYan, YUANChunqing, GAOXinyuan, et al. Analysis of disease characteristics and influencing factors of adolescent children with large bone age and short stature in Beijing from 2017 to 2022[J]. Chinese Journal of Health Statistics, 2023, 40(5): 744-747.
LIRuonan, SHANGXin, ZHANGTenglin, et al. Childhood obesity and central precocious puberty[J]. Journal of Central South University. Medical Science, 2024, 49(7): 1034-1041.
[32]
HuangJW, YangDM, RongRC, et al. A critical assessment of using ChatGPT for extracting structured data from clinical notes[J]. NPJ Digit Med, 2024, 7(1): 106.
[33]
WeiJ, WangXZ, SchuurmansD, et al. Chain-of-thought prompting elicits reasoning in large language models[EB/OL]. (2023-01-10)[2024-11-01].