Given the challenges of obtaining randomized controlled trial data, causal inference based on observational data has emerged as an alternative research paradigm. However, missing covariates in observational data often introduce bias, leading to distorted estimates of causal effects. To address this problem, this paper proposes a three-stage weighted causal effect estimation algorithm for handling missing covariates. First, covariates are categorized into confounders, exposure predictors, and outcome predictors based on the causal graph, and a multiple imputation strategy is implemented for missing confounders to preserve the integrity of the causal structure. Second, a missing pattern is introduced, and propensity score models are constructed through an improved covariate balancing strategy. Finally, causal effects are estimated by integrating inverse probability weighting across missing patterns. Simulation results demonstrate that, compared with existing methods such as gradient boosting machines, the proposed method reduces the root mean square error (RMSE) in most scenarios, validating its effectiveness in estimating causal effects under missing data conditions.
逻辑回归(Logistic Regression,LR)作为估计倾向得分的基准模型,在医学、社会科学、经济学等众多领域应用普遍,其原理清晰、结果易于解释。更复杂的加权策略,如模型平均、机器学习等方法可能会增加计算负担和结果的不确定性。逻辑回归为TSW方法提供了一个相对简洁、计算高效且易于与分组步骤结合的加权框架,便于验证该方法的有效性。在这种方法中,将处理变量T作为因变量,协变量 x 作为自变量,将LR输出结果作为倾向得分估计值。二元处理变量的倾向得分估计模型可以表示为:
FANJ Y, ZHANM F, CAIZ W, et al. Covariate balancing in propensity score estimation with variable selection: Based on GMM-LASSO approach[J]. Systems Engineering-Theory & Practice, 2021, 41(10):2631-2639. DOI:10.12011/SETP2020-0037(Ch ).
[3]
BOTTOUL, PETERSJ, QUIÑONERO-CANDELAJ, et al. Counterfactual reasoning and learning systems: The example of computational advertising[J] Journal of Machine Learning Research,2013, 14(1): 3207-3260. DOI:10.5555/2567709.2567766 .
[4]
ROSENBAUMP R, RUBIND B. The central role of the propensity score in observational studies for causal effects[J]. Biometrika, 1983, 70(1): 41-55. DOI: 10.1093/biomet/70.1.41 .
JIANGQ S, MAJ Y, HUANGC, et al. Covariate distribution balance via energy distance for causal inference[J]. Statistical Research, 2023, 40(5): 144-151. DOI:10.19343/j.cnki.11-1302/c.2023.05.011(Ch ).
[7]
ROBINSJ M, ROTNITZKYA, ZHAOL P. Estimation of regression coefficients when some regressors are not always observed[J]. Journal of the American Statistical Association, 1994, 89(427): 846-866. DOI:10.1080/01621459.1994.10476818 .
[8]
DUDÍKM, LANGFORDJ, LIL H. Doubly robust policy evaluation and learning[EB/OL].[2025-06-07]. DOI: 10.1214/14-sts500 .
[9]
RODRIGUEZ DUQUED, STEPHENSD A, MOODIEE E M, et al. Semiparametric Bayesian inference for optimal dynamic treatment regimes via dynamic marginal structural models[J]. Biostatistics, 2023, 24(3): 708-727. DOI:10.1093/biostatistics/kxac007 .
LIUC, SHAX K, ZHANGY. Staggered difference-in-differences method: Heterogeneous treatment effects and choice of estimation[J]. Journal of Quantitative & Technological Economics, 2022, 39(9): 177-204. DOI:10.13653/j.cnki.jqte.20220805.001(Ch ).
[12]
ANGRISTJ D, IMBENSG W, RUBIND B. Identification of causal effects using instrumental variables[J]. Journal of the American Statistical Association, 1996, 91(434): 444-455. DOI:10.1080/01621459.1996.10476902 .
[13]
ANDERSONM, MAGRUDERJ. Learning from the crowd: Regression discontinuity estimates of the effects of an online review database[J]. The Economic Journal, 2012, 122(563): 957-989. DOI:10.1111/j.1468-0297.2012.02512.x .
[14]
MIAOW, TCHETGEN TCHETGENE. Invited commentary: Bias attenuation and identification of causal effects with multiple negative controls[J]. American Journal of Epidemiology, 2017, 185(10): 950-953. DOI:10.1093/aje/kwx012 .
[15]
LIH K, MIAOW, CAIZ, et al. Causal data fusion methods using summary-level statistics for a continuous outcome[J]. Statistics in Medicine, 2020, 39(8): 1054-1067. DOI:10.1002/sim.8461 .
[16]
JOSEYK P, YANGF, GHOSHD, et al. A calibration approach to transportability and data-fusion with observational data[J]. Statistics in Medicine, 2022, 41(23): 4511-4531. DOI:10.1002/sim.9523 .
[17]
YANGS, DINGP. Combining multiple observational data sources to estimate causal effects[J]. Journal of the American Statistical Association, 2020, 115(531): 1540-1554. DOI:10.1080/01621459.2019.1609973 .
[18]
WHITEI R, CARLINJ B. Bias and efficiency of multiple imputation compared with complete-case analysis for missing covariate values[J]. Statistics in Medicine, 2010, 29(28): 2920-2931. DOI:10.1002/sim.3944 .
[19]
ROSENBAUMP R, RUBIND B. Reducing bias in observational studies using subclassification on the propensity score[J]. Journal of the American Statistical Association, 1984, 79(387): 516-524. DOI:10.1080/01621459.1984.10478078 .
[20]
D’AGOSTINOR, LANGW, WALKUPM, et al. Examining the impact of missing data on propensity score estimation in determining the effectiveness of self-monitoring of blood glucose (SMBG)[J]. Health Services and Outcomes Research Methodology, 2001, 2(3): 291-315. DOI:10.1023/A: 1020375413191 .
[21]
RUBIND B. Multiple Imputation for Nonresponse in Surveys[M]. Hoboken: Wiley, 1987. DOI:10.1002/9780470316696 .
DENGJ X, SHANL B, HED Q, et al. Processing method of missing data and its developing tendency[J]. Statistics & Decision, 2019, 35(23): 28-34. DOI:10.13546/j.cnki.tjyjc.2019.23.005(Ch ).
[24]
GRAHAMJ W. Missing data analysis: Making it work in the real world[J]. Annual Review of Psychology, 2009, 60: 549-576. DOI:10.1146/annurev.psych.58.110405.085530 .
[25]
MCCAFFREYD F, GRIFFINB A, ALMIRALLD, et al. A tutorial on propensity score estimation for multiple treatments using generalized boosted models[J]. Statistics in Medicine, 2013, 32(19): 3388-3414. DOI:10.1002/sim.5753 .
[26]
WANGZ Q, TANGN S. Bayesian quantile regression with mixed discrete and nonignorable missing covariates[J]. Bayesian Analysis, 2020, 15(2): 579-604. DOI:10.1214/19-ba1165 .
[27]
RUBIND B. Estimating causal effects of treatments in randomized and nonrandomized studies[J]. Journal of Educational Psychology, 1974, 66(5): 688-701. DOI:10.1037/h0037350 .
[28]
CROWEB J, LIPKOVICHI A, WANGO H. Comparison of several imputation methods for missing baseline data in propensity scores analysis of binary outcome[J]. Pharmaceutical Statistics, 2010, 9(4): 269-279. DOI:10.1002/pst.389 .
ZHUJ P, ZHENGC L, FANGK N. Two-stage credit scoring model with missing data: Based on the analysis of Internet consumer credit data[J]. Journal of Applied Statistics and Management, 2021, 40(4): 613-624. DOI:10.13860/j.cnki.sltj.20201219-006(Ch ).
[31]
SETOGUCHIS, SCHNEEWEISSS, BROOKHARTM A, et al. Evaluating uses of data mining techniques in propensity score estimation: A simulation study[J]. Pharmacoepidemiology and Drug Safety, 2008, 17(6): 546-555. DOI:10.1002/pds.1555 .