This paper concerns a scalable model⁃averaging method with right⁃censored responses. By using inverse probability censoring weighting to synthesize new responses and a singular value decomposition to transform the original models, this method enables us to estimate the weights by considering maximum of p covariate models. The Mallows criterion and the Jackknife criterion are applied to the selection of weights without the standard constraint that the weights sum to one. The theoretical results show that the calculated weights exhibit asymptotic optimality in the sense of achieving the lowest possible least squares error. Numerical studies demonstrate the excellent averaging performance of the proposed model in terms of both predictive accuracy and computational time.
模型平均是一种处理模型不确定性和提高预测精度的有效方法。该方法将一系列候选模型用适当的权重组合在一起,将每个候选模型中包含的信息纳入模型预测,从而得到更加稳健和精确的估计。模型平均主要分为贝叶斯模型平均和频率模型平均两个方向,其中基于渐近最优性的各种频率模型平均方法受到越来越多的关注和研究,如基于Mallows准则的Mallows model averaging(MMA)方法[1-2]、基于Jackknife准则的Jackknife model averaging(JMA)方法[3-4]、异方差数据下稳健的方法[5]、高维数据下模型平均方法[6]等。这些方法大多考虑完整数据的情形,但是在实际问题中经常会遇到不完整数据。如在生存分析中,生存时间经常会因各种原因出现右删失现象。针对删失数据,在局部误设定的框架下,文献[7-9]分别考虑了线性模型、分位数回归模型和部分线性分位数回归模型的模型平均问题。在去掉局部误设定的约束后,当响应变量右删失时,Yan等[10]研究了高维数据下的模型平均,Liang等[11]、Dong等[12]和Hu等[13]分别研究了线性模型及部分线性模型的模型平均估计,他们所建立方法都具有渐近最优性,且有很好的预测性能。
文献[11]采用Mallows准则[5]选择权重, 由此得到的right censored linear model averaging(RCLMA)估计是渐近最优的,且有很好的预测性能。但是,当协变量的个数p较大时,这种基于原模型的模型平均方法可能需要估计个权重,极大地降低了计算速度和预测精度。为了解决候选模型数量较大带来的计算困难,我们尝试采用新的方法去改进。参考文献[14],引入奇异值分解将协变量矩阵转换成列正交的协变量矩阵,转换后候选模型的数量降为p个,且候选模型均为简单的一元回归模型,从而有效地降低了计算压力和时间成本。
本节将通过模拟实验和实例分析考察所提的RC⁃SMMA、RC⁃SJMA方法在有限样本集上的表现,并与模型平均方法RCLMA[11]、weighted least squares model averaging(WLSMA)[12]、smoothed Akaike information criterion(SAIC)、smoothed Bayesian information criterion (SBIC)、等权重(equal weight,EW)方法,以及模型选择方法Akaike information criterion(AIC)、Bayesian information criterion (BIC)进行比较。
SUNZ M, MAJ Y, SUZ. FIC model selection and model averaging for linear model with censored response [J]. Scientia Sinica: Mathematica, 2013,43(7): 647-661. (in Chinese)
[9]
DUJ, ZHANGZ Z, XIET F. Focused information criterion and model averaging in censored quantile regression[J]. Metrika, 2017, 80: 547-570.
[10]
SUNZ M, SUNL Q, LUX L, et al. Frequentist model averaging estimation for the censored partial linear quantile regression model[J]. Journal of Statistical Planning and Inference, 2017, 189: 1-15.
[11]
YANX D, WANGH N, WANGW, et al. Optimal model averaging forecasting in high⁃dimensional survival analysis[J]. International Journal of Forecasting, 2021, 37(3): 1147-1155.
[12]
LIANGZ Q, CHENX L, ZHOUY Q. Mallows model averaging estimation for linear regression model with right censored data[J]. Acta Mathematicae Applicatae Sinica, English Series, 2022, 38(1): 5-23.
[13]
DONGQ K, LIUB X, ZHAOH. Weighted least squares model averaging for accelerated failure time models[J]. Computational Statistics and Data Analysis, 2023, 184: 107743.
[14]
HUG Z, CHENGW H, ZENGJ. Optimal model averaging for semiparametric partially linear models with censored data[J]. Mathematics, 2023, 11(3): 734.
[15]
ZHUR, WANGH Y, ZHANGX Y, et al. A scalable frequentist model averaging method[J]. Journal of Business & Economic Statistics, 2023, 41(4): 1228-1237.
[16]
CLAESKENSG, HJORTN L. Model selection and model averaging[M]. Cambridge: Cambridge University Press, 2008.
[17]
HUANGJ, MAS G, XIEH L. Regularized estimation in the accelerated failure time model with high⁃dimensional covariates[J]. Biometrics, 2006, 62(3): 813-820.