大语言模型在智能病历内涵质控中的应用研究

杨浩 ,  徐旻 ,  彭荷玲 ,  欧阳洋 ,  罗洁 ,  李红霞 ,  郑玲

西安交通大学学报(医学版) ›› 2026, Vol. 47 ›› Issue (3) : 477 -485.

PDF (8510KB)
西安交通大学学报(医学版) ›› 2026, Vol. 47 ›› Issue (3) : 477 -485. DOI: 10.7652/jdyxb202603011
智慧医疗专题

大语言模型在智能病历内涵质控中的应用研究

作者信息 +

Application of large language models in intelligent intrinsic quality control of electronic medical records

Author information +
文章历史 +
PDF (8713K)

摘要

目的 探索基于大语言模型(large language model, LLM)的质控智能体在电子病历(electronic medical record, EMR)内涵质控中的应用效果,以提升病历文书的合规性、完整性及跨文书逻辑一致性,为医院智慧管理提供解决方案。方法 依托机器学习及医疗大模型技术,构建集病历理解、辅助诊断与质控能力于一体的人工智能(AI)质控系统。系统采用数据层-算法层-能力层的三级分层架构,融合多源异构医学数据,运用自回归无监督预训练、Few-shot提示、思维链(chain-of-thought, CoT)推理及低秩自适应(low-rank adaptation, LoRA)微调技术,完成病历内涵质控场景模型的预训练与微调。并通过随机抽取5 000份临床病历样本,对比AI质控智能体在病历的完整性、规范性、逻辑性与符合性指标的效能,同时采用Fisher's exact检验评估AI质控应用后病历质量的改善,并测算AI智能体的准确率、召回率及F1-score。结果 在一致性验证中,AI质控智能体在门诊与住院病历场景下分别达到93.7%与96.1%的精确率,91.9%与94.8%的F1-score,总体一致率分别为92.4%与95.1%。应用AI质控系统后,4项质控指标错误率均显著下降:完整性错误率从21%降至5%(改善率76.2%,P=0.001)、规范性错误率从22%降至8%(改善率63.6%,P=0.009)、逻辑性错误率从13%降至3%(改善率76.9%,P=0.016)、符合性错误率从10%降至2%(改善率80.0%,P=0.033)。典型案例分析表明,系统在跨文档逻辑矛盾检测、治疗时序逻辑矛盾及罕见病合规性校验等复杂场景下表现出良好的推理与纠错能力。结论 基于LLM的AI质控系统能够有效提升EMR质控的准确性与效率,显著改善病历文书质量。未来结合多模态数据融合及多中心验证,有望进一步拓展系统在智慧医院建设和医疗质量管理数字化转型中的应用价值。

Abstract

Objective To explore the application of an intelligent quality control (QC) system based on large language models (LLMs) in the intrinsic quality control of electronic medical records (EMRs), with the goal of enhancing the compliance, completeness, and cross-document logical consistency of clinical documentation, thereby supporting intelligent hospital management. Methods Leveraging machine learning and medical large language model technologies, an intelligent QC system was developed that integrates EMR comprehension, diagnostic assistance, and quality control capabilities. The system adopted a three-tier vertical architecture of data layer-algorithm layer-capability layer, incorporating heterogeneous multi-source medical data and employing autoregressive unsupervised pre-training, few-shot prompting, chain-of-thought (CoT) reasoning, and low-rank adaptation (LoRA) fine-tuning techniques to complete pre-training and fine-tuning of the model for EMR intrinsic quality control scenarios. A random sample of 5 000 clinical medical records was selected to compare the performance of the AI QC agent across completeness, standardization, logical consistency, and conformity metrics. Fisher's exact test was used to evaluate the statistical significance of EMR quality improvement after AI quality control implementation; the accuracy, recall, and F1-score of the AI agent were calculated. Results In consistency validation, the AI QC agent achieved precision rates of 93.7% and 96.1% for outpatient and inpatient medical record scenarios respectively, with F1-scores of 91.9% and 94.8%, and overall consistency rates of 92.4% and 95.1%, respectively. After the intelligent QC system was implemented, error rates across all the four quality control indicators decreased significantly: completeness error rate decreased from 21% to 5% (76.2% improvement, P=0.001), standardization error rate decreased from 22% to 8% (63.6% improvement, P=0.009), logical consistency error rate decreased from 13% to 3% (76.9% improvement, P=0.016), and conformity error rate decreased from 10% to 2% (80.0% improvement, P=0.033). Typical case analysis demonstrated that the system exhibited excellent reasoning and error correction capabilities in complex scenarios such as cross-document logical contradiction detection, treatment temporal logic contradictions, and rare disease compliance verification. Conclusion The intelligent QC system based on LLMs effectively improves the accuracy and efficiency of EMR quality control and significantly enhances clinical documentation quality. It shows strong clinical applicability and scalability. Future work will focus on expanding the system's capabilities through multimodal data integration and multicenter validation, so as to further advance its role in smart hospital development and the digital transformation of healthcare quality management.

关键词

电子病历(EMR) / 大型语言模型(LLM) / 自动化质控 / 诊断编码准确性 / 医疗记录管理

Key words

electronic medical record (EMR) / large language model (LLM) / automated quality control / diagnostic coding accuracy / medical record management

引用本文

引用格式 ▾
杨浩,徐旻,彭荷玲,欧阳洋,罗洁,李红霞,郑玲. 大语言模型在智能病历内涵质控中的应用研究[J]. 西安交通大学学报(医学版), 2026, 47(3): 477-485 DOI:10.7652/jdyxb202603011

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

MA L, MA J P, LI X Y, et al. Application of hierarchical embedded large model architecture in governance of multimodal health and medical data[J]. Infor Technol Infor, 2025(1): 137—139, 146.

[2]

JIANG S Y, LI Y C, XUE W D, et al. Research on quality control method of multi—stage discharge summary based on large model[J]. Chin J Health Inform Manag, 2023, 20(6): 888-896.

[3]

SINGHAL K, AZIZI S, TU T, et al. Large language models encode clinical knowledge[J]. Nature, 2023, 620(7972): 172-180.

[4]

ZHOU C, LI Q, LI C, et al. A comprehensive survey on pre—trained foundation models: a history from BERT to ChatGPT[J]. Int J Mach Learn Cybern, 2024: 1-65.

[5]

THIRUNAVUKARASU A J, TING D S J, ELANGOVAN K, et al. Large language models in medicine[J]. Nat Med, 2023, 29(8): 1930-1940.

[6]

RUAN T, BIAN Y A, YU G Y, et al. A review on research and application of medical large language models[J]. Chin J Health Inform Manag, 2023, 20(6): 853-861.

[7]

PAN S D, XIANG P, WANG G X, et al. Assisted generation and quality control of the inpatient medical record based on MedCopilot[J]. Chin J Health Inform Manag, 2025, 22(1): 26-31, 44.

[8]

BAO Z J, CHEN W, XIAO S Z, et al. Disc—medllm: bridging general large language models and real—world medical consultation[DB/OL].arXiv, 2023: 14346. (2023—08—28) [2025—08—10].https://arxiv.org/abs/2308.14346.

[9]

PREIKSAITIS C, ROSE C. Opportunities, challenges, and future directions of generative artificial intelligence in medical education: scoping review[J]. JMIR Med Educ, 2023, 9: e48785.

[10]

BAI P F, HUANG Z H, WANG Y. A review on application and research of large language models in the smart hospitals[J]. Comput Appl Softw, 2024, 41(7): 1-5, 19.

[11]

MA W R, GONG M C, DAI H, et al. A comprehensive review of the applications of large language models in clinical medicine with ChatGPT as a representative[J]. J Med Inform, 2023, 44(7): 9-17.

[12]

MENG L G, WANG Y, LIU D, et al. Application of intelligent information technology in the field of medical equipment management[J]. J Xi'an Jiaotong Univ (Med Sci), 2024, 45(3): 520-524.

[13]

WEI J, WANG X Z, SCHUURMANS D, et al. Chain—of—thought prompting elicits reasoning in large language models[J]. Adv Neural Inf Process Syst, 2022, 35: 24824-24837.

[14]

LIEVIN V, HOTHER C E, MOTZFELDT A G, et al. Can large language models reason about medical questions?[J]. Patterns, 2024, 5(3): 100943.

[15]

HU E J, SHEN Y, WALLIS P, et al. Lora: low—rank adaptation of large language models[J]. ICLR, 2022, 1(2): 3.

[16]

WU C, LIN W, ZHANG X, et al. PMC—LLaMA: toward building open—source language models for medicine[J]. J Am Med Inform Assoc, 2024, 31(9): 1833-1843.

基金资助

国家肿瘤临床医学研究中心中青年研究基金资助(DSS-YSF-2023007)

国家卫生健康委医院管理研究所2023年医疗质量(循证)管理研究项目(YLZLXZ23G096)

国家卫生健康委医院管理研究所2024年医疗人工智能临床应用研究(YLXX24AIA045)

AI Summary AI Mindmap
PDF (8510KB)

198

访问

0

被引

详细

导航
相关文章

AI思维导图

/