大语言模型预测手术持续时长对手术资源配置管理效率的促进作用

王毅豪 ,  袁骏毅 ,  张蕾 ,  舒婷

西安交通大学学报(医学版) ›› 2026, Vol. 47 ›› Issue (3) : 464 -470.

PDF (3303KB)
西安交通大学学报(医学版) ›› 2026, Vol. 47 ›› Issue (3) : 464 -470. DOI: 10.7652/jdyxb202603009
智慧医疗专题

大语言模型预测手术持续时长对手术资源配置管理效率的促进作用

作者信息 +

Large language models predict the promoting effect of operation duration on the efficiency of surgical resource allocation and management

Author information +
文章历史 +
PDF (3381K)

摘要

目的 通过比较不同大语言模型对手术持续时长的预测性能,评估其提高手术室资源利用效率的潜在价值。方法 选取2024年11月至2025年1月上海市胸科医院的6 154条外科手术数据(手术名称、主刀医生、一助医生等)进行大语言模型(Qwen-7B、DeepSeek-32B和 DeepSeek-671B)的训练和测试,主要采用 LoRA 微调技术、检索增强生成策略以及提示词工程等方法对模型进行优化,并通过均方误差、平均绝对误差、平均绝对百分比误差以及手术持续时长预测准确率等指标评估模型性能。此外,邀请3位医院管理者在大语言模型辅助下进行手术排程,通过与未使用大语言模型辅助时的手术室占用时间进行比较,评估大语言模型对手术资源配置管理的辅助效果。结果 DeepSeek-671B模型的预测准确率为52.06%(Z=-6.695,P<0.001),显著高于 Qwen-7B 的45.99%(Z=-2.854,P<0.001)和DeepSeek-32B的48.16%(Z=-5.199,P<0.005);同时,DeepSeek-671B的回归误差指标也优于 Qwen-7B和 DeepSeek-32B(MSE:2 486.11 vs.3 734.31 vs.3 224.89,MAE:34.99 vs.40.31 vs.38.78);三种大语言模型辅助前后,手术室占用时间分别减少了5.48%、3.37%、8.26%,秩和检验结果显示具有统计学意义(Z= -3.408,P<0.005)。结论 通过大语言模型预测手术持续时长有助于提升手术资源配置的管理效率,医院管理者能够更科学地安排手术顺序,从而有效提升手术室整体运行效能。

Abstract

Objective To evaluate the potential value of large language models in improving the efficiency of operating room resource utilization by comparing the performance of different large language models in predicting the duration of surgery. Methods A total of 6 154 surgical operation data (mainly including operation name, chief surgeon, and assistant doctor) from Shanghai Chest Hospital from November 2024 to January 2025 were selected for large language models (Qwen-7B, DeepSeek-32B and DeepSeek-671B) training and testing. The LoRA fine-tuning technology, retrieval enhancement generation strategy, and cue word engineering were used to optimize the model; the performance of the model was evaluated by the mean square error, mean absolute error, mean absolute percentage error, and the accuracy of surgery duration prediction. In addition, three hospital managers were invited to perform surgery scheduling with the assistance of large language models, and the auxiliary effect of large language models on surgical resource allocation management was evaluated by comparing the occupancy time of the operating room with that without the assistance of large language models. Results The prediction accuracy of DeepSeek-671B model was 52.06% (Z=-6.695, P<0.001), which was significantly higher than that of Qwen-7B 45.99% (Z=-2.854, P<0.001) and DeepSeek-32B 48.16% (Z=-5.199, P<0.001). Meanwhile, the regression error index of DeepSeek-671B was also better than that of Qwen-7B and DeepSeek-32B (MSE: 2 486.11 vs.3 734.31 vs. 3 224.89, MAE: 34.99 vs. 40.31 vs. 38.78). Before and after the assistance of the three large language models, the actual operating room occupancy time was reduced by 5.48%, 3.37% and 8.26%, respectively, and the rank sum test results showed that the difference was statistically significant (Z= -3.408, P<0.005). Conclusion Predicting the duration of surgery by large language models helps to improve the management efficiency of surgical resource allocation, and hospital managers can arrange the operation sequence more scientifically so as to effectively improve the overall operation efficiency of the operating room.

关键词

医院管理 / 手术资源配置 / 大语言模型 / Qwen / DeepSeek / 模型微调 / 检索增强生成

Key words

hospital management / surgical resources allocation / large language model / Qwen / DeepSeek / model fine-tuning / retrieval augmented generation

引用本文

引用格式 ▾
王毅豪,袁骏毅,张蕾,舒婷. 大语言模型预测手术持续时长对手术资源配置管理效率的促进作用[J]. 西安交通大学学报(医学版), 2026, 47(3): 464-470 DOI:10.7652/jdyxb202603009

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

雷甜甜, 梁鹏, 马洪升, . 四川大学华西医院日归手术管理实践[J]. 广东医学, 2022, 43(10): 1222-1228.

[2]

LEI T T, LIANG P, MA H S, et al. Management practice of day return surgery in West China Hospital of Sichuan University[J]. Guangdong Med J, 2022, 43(10): 1222-1228.

[3]

刘静, 陈红, 吴波, . 移动手术护理排班系统的设计与应用研究[J]. 中国数字医学, 2022, 17(8): 61-65.

[4]

LIU J, CHEN H, WU B, et al. Design and application of mobile surgical nursing scheduling system[J]. China Digit Med, 2022, 17(8): 61-65.

[5]

姚曼琳, 王筱君, 郝雪梅, . 基于互联网的医院手术室排班信息平台的构建及应用[J]. 中国医学装备, 2023, 20(3): 122-125.

[6]

YAO M L, WANG X J, HAO X M, et al. The construction of the internet—based scheduling informatization platform of operating room of hospital and the application of that in operating room[J]. China Med Equip, 2023, 20(3): 122-125.

[7]

LEVIN J M, ZARIBAFZADEH H, DOYLE T R, et al. A machine learning prediction model for total shoulder arthroplasty procedure duration: an evaluation of surgeon, patient, and shoulder—specific factors[J]. J Shoulder Elbow Surg, 2025, 34(7): 1792-1800.

[8]

RAMAMURTHI A, NEUPANE B, DESHPANDE P, et al. Development and validation of an artificial intelligence system for surgical case length prediction[J]. Surgery, 2025, 179: 108942.

[9]

HE J, RUNGTA M, KOLECZEK D, et al. Does prompt formatting have any impact on LLM performance?[DB/OL].arXiv, 2024: 2411.10541. (2024—11—15)[2025—03—22].https://arxiv.org/abs/2411.10541.

[10]

ZHANG B, LIU Z T, CHERRY C, et al. When scaling meets LLM finetuning: the effect of data, model and finetuning method[DB/OL].arXiv, 2024: 17193. (2024—02—27) [2025—03—22].https://arxiv.org/abs/2402.17193.

[11]

HU J C, LIAO X X, GAO J, et al. Optimizing large language models with an enhanced LoRA fine—tuning algorithm for efficiency and robustness in NLP tasks[DB/OL].arXiv, 2024: 18729. (2024—12—25) [2025—03—22].https://arxiv.org/abs/2412.18729.

[12]

MIAO J, THONGPRAYOON C, SUPPADUNGSUK S, et al. Integrating retrieval—augmented generation with large language models in nephrology: advancing practical applications[J]. Medicina (Kaunas), 2024, 60(3): 445.

[13]

GAO Y F, XIONG Y, GAO X Y, et al. Retrieval—augmented generation for large language models: a survey[DB/OL].arXiv, 2023: 10997. (2024—03—27) [2025—03—22].https://arxiv.org/abs/2312.10997.

[14]

MIAO J, THONGPRAYOON C, SUPPADUNGSUK S, et al. Chain of thought utilization in large language models and application in nephrology[J]. Medicina (Kaunas), 2024, 60(1): 148.

[15]

LI J H, XU J M, HUANG S, et al. Large language model inference acceleration: a comprehensive hardware perspective[DB/OL].arXiv, 2024: 04466. (2025—06—13) [2025—03—22].https://arxiv.org/abs/2410.04466.

[16]

ZHOU Z X, NING X F, HONG K, et al. A survey on efficient inference for large language models[DB/OL].arXiv, 2024: 14294. (2024—05—22) [2025—03—22].https://arxiv.org/abs/2404.14294.

[17]

袁骏毅, 戴锦杰 . 基于数字孪生技术的日间手术管理平台构建及应用[J]. 中国卫生质量管理, 2023, 30(8): 35-38.

[18]

YUAN J Y, DAI J J. Construction and application of ambulatory surgery management platform driven by digital twin[J]. Chin Health Qual Manag, 2023, 30(8): 35-38.

[19]

HASAN M R, RAY R K, CHOWDHURY F R. Employee performance prediction: an integrated approach of business analytics and machine learning[J]. J Manage Stud, 2024, 6(1): 215-219.

[20]

MATHIVANAN S K, SONAIMUTHU S, MURUGESAN S, et al. Employing deep learning and transfer learning for accurate brain tumor detection[J]. Sci Rep, 2024, 14(1): 7232.

基金资助

国家卫生健康委医院管理研究所医疗人工智能临床应用研究项目(YLXX24AIC003)

上海申康医院发展中心技术规范化管理和推广项目(SHDC22026202)

上海市卫生健康委员会智慧医疗专项研究项目(2025ZHYL011)

AI Summary AI Mindmap
PDF (3303KB)

43

访问

0

被引

详细

导航
相关文章

AI思维导图

/