采用机器学习的大规模并行程序I/O性能建模与预测方法

刘恒 ,  王子衡 ,  王强 ,  刘新栋 ,  龙炳志 ,  董小社

西安交通大学学报 ›› 2026, Vol. 60 ›› Issue (7) : 207 -218.

PDF (4389KB)
西安交通大学学报 ›› 2026, Vol. 60 ›› Issue (7) : 207 -218. DOI: 10.7652/xjtuxb202607019
专题 高纯气体制备

采用机器学习的大规模并行程序I/O性能建模与预测方法

作者信息 +

I/O Performance Modeling and Prediction Method for Large-Scale Parallel Applications Using Machine Learning

Author information +
文章历史 +
PDF (4493K)

摘要

针对大规模并行程序因输入/输出(I/O)数据难以获取、建模与优化成本高昂而导致的性能分析效率低下和调优困难问题,提出了一种新的机器学习驱动的大规模并行程序I/O性能建模与预测方法。该方法在建模阶段,一方面基于小规模节点环境下的采样数据,利用线性回归方法构建可用于大规模节点外推的并行程序I/O特征预测模型;另一方面,将小规模节点环境下采集的并行程序I/O特征与对应的I/O栈参数空间共同输入人工神经网络(ANN)进行训练,以学习系统配置参数与I/O性能之间的非线性映射关系,从而构建并行程序的I/O性能预测模型。在预测阶段,使用I/O特征预测模型预测大规模并行程序的I/O特征,并将其与对应的大规模并行程序I/O栈参数空间输入I/O性能预测模型,实现对大规模并行程序I/O性能的准确外推预测。实验结果表明:在国产超级计算机上,采用所提方法在4种典型测试程序IOR、S3D-IO、BT-IO和Flash-IO上的预测精度均较高;当采用1~16节点规模下训练得到的模型对128节点(2 048进程)场景进行外推预测时,其平均绝对百分比误差分别为19.07%、18.96%、12.34%和14.16%;基于所构建的I/O性能预测模型对I/O栈参数调优后,4种典型程序I/O性能加速比分别达到16.38、23.16、45.36和65.38倍。

Abstract

To address the issues of inefficient performance analysis and difficulty in tuning caused by the scarcity of input/output (I/O) data and the prohibitive costs of modeling and optimization in large-scale parallel applications, a novel method for machine learning-driven I/O performance modeling and prediction in large-scale parallel applications is proposed. During the modeling phase, on the one hand, a linear regression-based model for I/O feature prediction in parallel applications is constructed using data sampled from small-scale node environments to enable the extrapolation of features to large-scale configurations; on the other hand, I/O features in parallel applications acquired from small-scale node environments, along with the corresponding I/O stack parameter space, are fed into an artificial neural network (ANN) for training. This allows for the learning of the nonlinear mapping between system configuration parameters and I/O performance, leading to the construction of a model for I/O performance prediction in parallel applications. During the prediction phase, the I/O features of large-scale parallel applications are first predicted via the I/O feature prediction model. These predicted features, together with the corresponding I/O stack parameter space of large-scale parallel applications, are then fed into the I/O performance prediction model to accurately predict the I/O performance of such applications. Experimental results demonstrate that, on a domestic supercomputer, the proposed method exhibits high prediction accuracy across four benchmark applications: IOR, S3D-IO, BT-IO, and Flash-IO. When models trained on 1—16 nodes are extrapolated to a 128-node (2 048-process) scenario, the mean absolute percentage errors (MAPEs) are 19.07%, 18.96%, 12.34%, and 14.16%, respectively. Furthermore, by tuning the I/O stack parameters based on the proposed model, I/O performance speedups of 16.38, 23.16, 45.36, and 65.38-fold are achieved for the four benchmark applications.

关键词

大规模并行程序 / I/O性能建模 / I/O自动调优 / 人工神经网络

Key words

large-scale parallel applications / I/O performance modeling / I/O auto-tuning / artificial neural network

引用本文

引用格式 ▾
刘恒,王子衡,王强,刘新栋,龙炳志,董小社. 采用机器学习的大规模并行程序I/O性能建模与预测方法[J]. 西安交通大学学报, 2026, 60(7): 207-218 DOI:10.7652/xjtuxb202607019

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

Luu H, Winslett M, Gropp W, et al. A multiplatform study of I/O behavior on petascale supercomputers[C]//Proceedings of the 24th International Symposium on High—Performance Parallel and Distributed Computing. New York, USA: ACM, 2015: 33-44.

[2]

张成 . 基于回归分析和集成学习的HPC应用I/O性能优化方法研究[D].西安: 西北大学, 2023.

[3]

孙经纬 . 数据驱动的高性能计算程序执行时间预测与优化研究[D].合肥: 中国科学技术大学, 2020.

[4]

Zhu Zhaobin, Neuwirth S, Lippert T . A comprehensive I/O knowledge cycle for modular and automated HPC workload analysis[C]//2022 IEEE International Conference on Cluster Computing (CLUSTER). Piscataway, NJ, USA: IEEE,2022: 581-588.

[5]

Liu Wei, Wu Kai, Liu Jialin, et al. Performance evaluation and modeling of HPC I/O on non—volatile memory[C]//2017 International Conference on Networking, Architecture, and Storage (NAS). Piscataway, NJ, USA: IEEE, 2017: 1-10.

[6]

Neuwirth S, Paul A K . Parallel I/O evaluation techniques and emerging HPC workloads: a perspective[C]//2021 IEEE International Conference on Cluster Computing (CLUSTER). Piscataway, NJ, USA: IEEE, 2021: 671-679.

[7]

汤志航, 兰颢, 刘政国, . HiTrain: 面向大模型训练的异构内存卸载与I/O优化[J].计算机研究与发展, 2026, 63(3): 627-639.

[8]

Tang Zhihang, Lan Hao, Liu Zhengguo, et al. HiTrain: heterogeneous memory offloading and I/O optimization for large language model training[J].Journal of Computer Research and Development, 2026, 63(3): 627-639.

[9]

程稳, 李焱, 曾令仿, . 面向Lustre集群存储的应用日志分析及系统自动优化框架[J].计算机工程与科学, 2022, 44(4): 594-604.

[10]

Cheng Wen, Li Yan, Zeng Lingfang, et al. An application log analysis and system automation optimization framework for Lustre cluster storage[J].Computer Engineering & Science, 2022, 44(4): 594-604.

[11]

张文韬, 汪璐, 程耀东 . 基于强化学习的Lustre文件系统的性能调优[J].计算机研究与发展, 2019, 56(7): 1578-1586.

[12]

Zhang Wentao, Wang Lu, Cheng Yaodong . Performance optimization of Lustre file system based on reinforcement learning[J].Journal of Computer Research and Development, 2019, 56(7): 1578-1586.

[13]

田鸿运, 武林平, 董勇, . 面向大规模集群的并行I/O用户层配置优化策略[J].国防科技大学学报, 2020, 42(2): 23-30.

[14]

Tian Hongyun, Wu Linping, Dong Yong, et al. User—level parallel I/O configuration optimize strategy toward large—scale cluster[J].Journal of National University of Defense Technology, 2020, 42(2): 23-30.

[15]

Egersdoerfer C, Rashid M H, Dai Dong, et al. Understanding and predicting cross—application I/O interference in HPC storage systems[C]//SC24—W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis. Piscataway, NJ, USA: IEEE, 2024: 1330-1339.

[16]

Behzad B, Luu H V T, Huchette J, et al. Taming parallel I/O complexity with auto—tuning[C]//Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis. New York, USA: ACM, 2013: 68.

[17]

Tipu A J S, Conbhuí P Ó, Howley E . Artificial neural networks based predictions towards the auto—tuning and optimization of parallel IO bandwidth in HPC system[J].Cluster Computing, 2024, 27(1): 71-90.

[18]

Nicolas L M, Mimouni S, Couvée P, et al. I/O patterns modeling of HPC applications with call stacks for predictive prefetch[J].Future Generation Computer Systems, 2026, 175: 108034.

[19]

Behzad B, Byna S, Prabhat, et al. Optimizing I/O performance of HPC applications with autotuning[J].ACM Transactions on Parallel Computing, 2019, 5(4): 15.

[20]

Bagbaba A, Wang Xuan . Improving the mpi—io performance of applications with genetic algorithm based auto—tuning[C]//2021 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). Piscataway, NJ, USA: IEEE, 2021: 798-805.

[21]

Nemirovsky D, Arkose T, Markovic N, et al. A machine learning approach for performance prediction and scheduling on heterogeneous CPUs[C]//2017 29th International Symposium on Computer Architecture and High Performance Computing (SBAC—PAD). Piscataway, NJ, USA: IEEE, 2017: 121-128.

[22]

Kim S, Sim A, WU Kesheng, et al. Design and implementation of I/O performance prediction scheme on HPC systems through large—scale log analysis[J].Journal of Big Data, 2023, 10(1): 65.

[23]

Liu Zhangyu, Zhang Cheng, Wu Huijun, et al. Optimizing HPC I/O performance with regression analysis and ensemble learning[C]//2023 IEEE International Conference on Cluster Computing (CLUSTER). Piscataway, NJ, USA: IEEE,2023: 234-246.

[24]

Meswani M R, Laurenzano M A, Carrington L, et al. Modeling and predicting disk I/O time of HPC applications[C]//2010 DoD High Performance Computing Modernization Program Users Group Conference. Piscataway, NJ, USA: IEEE, 2010: 478-486.

[25]

Wang Wanxin, Wu Huijun, Yang Lihua, et al. AIO: automating I/O optimization pipeline for data—intensive applications in HPC[C]//2024 IEEE International Symposium on Parallel and Distributed Processing with Applications (ISPA). Piscataway, NJ, USA: IEEE, 2024: 1541-1548.

[26]

Shan Hongzhang, Shalf J. Using IOR to analyze the I/O performance for HPC platforms[EB/OL]. [2015—10—22].https://escholarship.org/uc/item/9111c60j.

[27]

Mendez S, Rexachs D, Luque E . Analyzing the parallel I/O severity of MPI applications[C]//2017 17th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID). Piscataway, NJ, USA: IEEE, 2017: 953-962.

[28]

Agarwal M, Singhvi D, Malakar P, et al. Active learning—based automatic tuning and prediction of parallel I/O performance[C]//2019 IEEE/ACM Fourth International Parallel Data Systems Workshop (PDSW). Piscataway, NJ, USA: IEEE, 2019: 20-29.

[29]

Rajesh N, Bateman K, Bez J L, et al. TunIO: an AI—powered framework for optimizing HPC I/O[C]//2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS). Piscataway, NJ, USA: IEEE, 2024: 494-505.

[30]

Chen Si, De Gonzalo S G, Wildani A . Few—shot HPC application runtime prediction[C]//2023 IEEE International Conference on Cluster Computing Workshops (CLUSTER Workshops). Piscataway, NJ, USA: IEEE,2023: 46-47.

基金资助

国家重点研发计划资助项目(2023YFB3001804)

AI Summary AI Mindmap
PDF (4389KB)

0

访问

0

被引

详细

导航
相关文章

AI思维导图

/