PDF (1398K)
摘要
共轭梯度方法(CG)和稳定双共轭梯度方法(BiCGSTAB)是求解稀疏线性系统的两种经典且高效的迭代方法,被广泛应用于科学计算和工程问题中。尽管GPU等并行处理器提升了这两种方法的并行性,但最新的硬件单元Tensor Core及其计算能力尚未被用于这两种方法中。该文设计了一个Tensor Core加速的CG解法器,利用Tensor Core计算CG和BiCGSTAB方法中的关键组件稀疏矩阵−向量乘法(SpMV)和点积操作,以发挥Tensor Core的计算能力,从而提升两种方法的整体性能。在NVIDIA A100和H100 GPU上的实验结果表明,Tensor Core加速的这两种方法相比调用CUDA官方库的基准版本在多个稀疏矩阵上均取得了显著的加速效果。
Abstract
Conjugate gradient (CG) and biconjugate gradient stabilized (BiCGSTAB) are two classical and efficient iterative methods for solving sparse linear systems, widely used in scientific computing and engineering applications. Although GPUs and other parallel processors have enhanced the parallelism of these methods, the latest hardware unit, Tensor Core, and its computing power have not yet been fully exploited for these two methods. This work proposes a Tensor Core-accelerated CG solver that leverages Tensor Cores for the key components in the CG and BiCGSTAB methods, such as sparse matrix-vector multiplication (SpMV) and dot product computation, thereby exploiting the computational capability of Tensor Cores to improve the overall performance of both methods. Experimental results on NVIDIA A100 and H100 GPUs demonstrate that both of these methods accelerated by Tensor Core achieve significant speedups over the baseline version that uses the CUDA official library on various sparse matrices.
关键词
Key words
[Author(id=1276507863066394766, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, orderNo=0, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=null, emailSecond=null, emailThird=null, correspondingAuthor=0, authorType=1, ext={EN=AuthorExt(id=1276507863137697936, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, authorId=1276507863066394766, language=EN, stringName=Yuechen LU, firstName=Yuechen, middleName=null, lastName=LU, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=College of Artificial Intelligence , China University of Petroleum— Beijing, Beijing 102249, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1276507863196418193, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, authorId=1276507863066394766, language=CN, stringName=卢玥辰, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=中国石油大学(北京) 人工智能学院 , 北京 102249, bio={"content":"卢玥辰,博士生,主要从事稀疏矩阵计算方面的研究。
"}, bioImg=null, bioContent=卢玥辰,博士生,主要从事稀疏矩阵计算方面的研究。
, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1276507862974120074, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, xref=null, ext=[AuthorCompanyExt(id=1276507862990897291, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, companyId=1276507862974120074, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=College of Artificial Intelligence , China University of Petroleum— Beijing, Beijing 102249, China), AuthorCompanyExt(id=1276507863007674508, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, companyId=1276507862974120074, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=中国石油大学(北京) 人工智能学院 , 北京 102249)])]), Author(id=1276507863250944147, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, orderNo=1, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=null, emailSecond=null, emailThird=null, correspondingAuthor=0, authorType=1, ext={EN=AuthorExt(id=1276507863326441621, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, authorId=1276507863250944147, language=EN, stringName=Yuxiao YUAN, firstName=Yuxiao, middleName=null, lastName=YUAN, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=College of Artificial Intelligence , China University of Petroleum— Beijing, Beijing 102249, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1276507863376773270, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, authorId=1276507863250944147, language=CN, stringName=袁雨萧, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=中国石油大学(北京) 人工智能学院 , 北京 102249, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1276507862974120074, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, xref=null, ext=[AuthorCompanyExt(id=1276507862990897291, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, companyId=1276507862974120074, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=College of Artificial Intelligence , China University of Petroleum— Beijing, Beijing 102249, China), AuthorCompanyExt(id=1276507863007674508, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, companyId=1276507862974120074, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=中国石油大学(北京) 人工智能学院 , 北京 102249)])]), Author(id=1276507863452270744, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, orderNo=2, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=null, emailSecond=null, emailThird=null, correspondingAuthor=0, authorType=1, ext={EN=AuthorExt(id=1276507863569711258, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, authorId=1276507863452270744, language=EN, stringName=Dechuang YANG, firstName=Dechuang, middleName=null, lastName=YANG, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=College of Artificial Intelligence , China University of Petroleum— Beijing, Beijing 102249, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1276507863628431515, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, authorId=1276507863452270744, language=CN, stringName=杨德闯, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=中国石油大学(北京) 人工智能学院 , 北京 102249, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1276507862974120074, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, xref=null, ext=[AuthorCompanyExt(id=1276507862990897291, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, companyId=1276507862974120074, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=College of Artificial Intelligence , China University of Petroleum— Beijing, Beijing 102249, China), AuthorCompanyExt(id=1276507863007674508, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, companyId=1276507862974120074, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=中国石油大学(北京) 人工智能学院 , 北京 102249)])]), Author(id=1276507863691346077, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, orderNo=3, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=weifeng.liu@cup.edu.cn, emailSecond=null, emailThird=null, correspondingAuthor=1, authorType=1, ext={EN=AuthorExt(id=1276507863779426463, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, authorId=1276507863691346077, language=EN, stringName=Weifeng LIU, firstName=Weifeng, middleName=null, lastName=LIU, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=*, address=College of Artificial Intelligence , China University of Petroleum— Beijing, Beijing 102249, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1276507863838146720, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, authorId=1276507863691346077, language=CN, stringName=刘伟峰, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=*, address=中国石油大学(北京) 人工智能学院 , 北京 102249, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1276507862974120074, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, xref=null, ext=[AuthorCompanyExt(id=1276507862990897291, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, companyId=1276507862974120074, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=College of Artificial Intelligence , China University of Petroleum— Beijing, Beijing 102249, China), AuthorCompanyExt(id=1276507863007674508, tenantId=1045748351789510663, journalId=1155139928303341607, articleId=1248992741033283960, companyId=1276507862974120074, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=中国石油大学(北京) 人工智能学院 , 北京 102249)])])]
卢玥辰,袁雨萧,杨德闯,刘伟峰.
GPU上
Tensor Core加速的共轭梯度解法器[J].
电子科技大学学报, 2026, 55(2): 244-251 DOI:10.12178/1001-0548.2024358
| [1] |
HESTENES M R, STIEFEL E. Methods of conjugate gradients for solving linear systems[J]. Journal of Research of the National Bureau of Standards, 1952, 49(6): 409-436.
|
| [2] |
VORST H A. Bi—CGSTAB: A fast and smoothly converging variant of Bi—CG for the solution of nonsymmetric linear systems[J]. SIAM Journal on Scientific and Statistical Computing, 1992, 13(2): 631-644.
|
| [3] |
李亿渊, 薛巍, 陈德训, 等. 稀疏矩阵向量乘法在申威众核架构上的性能优化[J]. 计算机学报, 2020, 43(6): 1010-1024.
|
| [4] |
LI Y Y, XUE W, CHEN D X, et al. Performance optimization for sparse matrix—vector multiplication on sunway architecture[J]. Chinese Journal of Computers, 2020, 43(6): 1010-1024.
|
| [5] |
吴洋, 赵永华, 纪国良 . 一类大规模稀疏矩阵特征问题求解的并行算法[J]. 数值计算与计算机应用, 2013, 34(2): 136-146.
|
| [6] |
WU Y, ZHAO Y H, JI G L. Parallel solving large—scale sparse matrix eigenvalue problem[J]. Journal on Numerical Methods and Computer Applications, 2013, 34(2): 136-146.
|
| [7] |
李佳佳, 张秀霞, 谭光明, 等. 选择稀疏矩阵乘法最优存储格式的研究[J]. 计算机研究与发展, 2014, 51(4): 882-894.
|
| [8] |
LI J J, ZHANG X X, TAN G M, et al. Study of choosing the optimal storage format of sparse matrix vector multiplication[J]. Journal of Computer Research and Development, 2014, 51(4): 882-894.
|
| [9] |
杜臻, 谭光明 . 稀疏矩阵向量乘的自动调优[J]. 计算物理, 2024, 41(1): 33-39.
|
| [10] |
DU Z, TAN G M. Auto—tuning for sparse matrix—vector multiplication[J]. Chinese Journal of Computational Physics, 2024, 41(1): 33-39.
|
| [11] |
徐小文 . 并行代数多重网格算法: 大规模计算应用现状与挑战[J]. 数值计算与计算机应用, 2019, 40(4): 243-260.
|
| [12] |
XU X W. Parallel algebraic multigrid methods: State—of—the art and challenges for extreme—scale applications[J]. Journal on Numerical Methods and Computer Applications, 2019, 40(4): 243-260.
|
| [13] |
毛润彰, 杜皓, 田鸿运, 等. 几类典型应用的代数多重网格算法并行可扩展瓶颈分析[J]. 计算物理, 2024, 41(4): 403-417.
|
| [14] |
MAO R Z, DU H, TIAN H Y, et al. Analysis of parallel scalability bottleneck for algebraic multigrid in typical real applications[J]. Chinese Journal of Computational Physics, 2024, 41(4): 403-417.
|
| [15] |
迟利华, 刘杰, 李晓梅 . 稀疏近似逆预条件子及其并行计算[J]. 计算机学报, 2000, 23(3): 255-260.
|
| [16] |
CHI L H, LIU J, LI X M. Parallel sparse approximate inverse preconditioners[J]. Chinese Journal of Computers, 2000, 23(3): 255-260.
|
| [17] |
GRIGORI L, TISSOT O. Scalable linear solvers based on enlarged Krylov subspaces with dynamic reduction of search directions[J]. SIAM Journal on Scientific Computing, 2019, 41(5): C522-C547.
|
| [18] |
YAMAZAKI I, CARSON E, KELLEY B. Mixed precision s—step conjugate gradient with residual replacement on GPUs[C]// Proceedings of the IEEE International Parallel and Distributed Processing Symposium. New York: IEEE, 2022: 886-896.
|
| [19] |
BURGESS J. RTX on—The NVIDIA turing GPU[C]// Proceedings of the IEEE Hot Chips 31 Symposium. New York: IEEE, 2019: 1-27.
|
| [20] |
CHOQUETTE J, GANDHI W. NVIDIA A100 GPU: Performance & innovation for GPU computing[C]// Proceedings of the IEEE Hot Chips 32 Symposium. New York: IEEE, 2020: 1-43.
|
| [21] |
CHOQUETTE J. Nvidia hopper GPU: Scaling performance[C]// Proceedings of the IEEE Hot Chips 34 Symposium. New York: IEEE, 2022: 1-46.
|
| [22] |
DAVIS T A, HU Y F. The university of Florida sparse matrix collection[J]. ACM Transactions on Mathematical Software, 2011, 38(1): 1-25.
|
| [23] |
LU Y C, LIU W F. DASP: Specific dense matrix multiply—accumulate units accelerated general sparse matrix—vector multiplication[C]// Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. New York: ACM, 2023: 1-14.
|
| [24] |
DAKKAK A, LI C, XIONG J J, et al. Accelerating reduction and scan using tensor core units[C]// Proceedings of the ACM International Conference on Supercomputing. New York: ACM, 2019: 46-57.
|
| [25] |
CHEN Y T, LI K, WANG Y H, et al. ConvStencil: Transform stencil computation to matrix multiplication on tensor cores[C]// Proceedings of the 29th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming. New York: ACM, 2024: 333-347.
|
| [26] |
OKANOVIC P, KWASNIEWSKI G, LABINI P S, et al. High performance unstructured SpMM computation using tensor cores[C]// Proceedings of the SC24: International Conference for High Performance Computing, Networking, Storage and Analysis. New York: IEEE, 2024: 1-14.
|
| [27] |
LU Y C, ZENG L J, WANG T C, et al. AmgT: Algebraic multigrid solver on tensor cores[C]// Proceedings of the SC24: International Conference for High Performance Computing, Networking, Storage and Analysis. New York: IEEE, 2024: 1-16.
|
基金资助
国家自然科学基金(U23A20301)
国家自然科学基金(62372467)