[论文解读] Cross-Fitting and Averaging for Machine Learning Estimation of Heterogeneous Treatment Effects
本文评估了在使用四种元学习器(T-learner、DR-learner、R-learner、X-learner)进行机器学习异质性处理效应估计时,样本分割、交叉拟合与平均策略的效果。研究发现,将五重交叉拟合与20次以上的迭代中位数平均相结合,可在复杂数据结构下实现最低的均方误差,并建议排除套索回归以降低方差并提高稳健性。
We investigate the finite sample performance of sample splitting, cross-fitting and averaging for the estimation of the conditional average treatment effect. Recently proposed methods, so-called meta-learners, make use of machine learning to estimate different nuisance functions and hence allow for fewer restrictions on the underlying structure of the data. To limit a potential overfitting bias that may result when using machine learning methods, cross-fitting estimators have been proposed. This includes the splitting of the data in different folds to reduce bias and averaging over folds to restore efficiency. To the best of our knowledge, it is not yet clear how exactly the data should be split and averaged. We employ a Monte Carlo study with different data generation processes and consider twelve different estimators that vary in sample-splitting, cross-fitting and averaging procedures. We investigate the performance of each estimator independently on four different meta-learners: the doubly-robust-learner, R-learner, T-learner and X-learner. We find that the performance of all meta-learners heavily depends on the procedure of splitting and averaging. The best performance in terms of mean squared error (MSE) among the sample split estimators can be achieved when applying cross-fitting plus taking the median over multiple different sample-splitting iterations. Some meta-learners exhibit a high variance when the lasso is included in the ML methods. Excluding the lasso decreases the variance and leads to robust and at least competitive results.
研究动机与目标
- 评估不同样本分割、交叉拟合与平均程序在估计异质性处理效应时的有限样本性能。
- 评估不同数据生成过程(随机对照试验与观察性研究)对元学习器估计量性能的影响。
- 确定在高维或非线性设定下,能最小化偏差与均方误差(MSE)的最优分割与平均策略。
- 探究在机器学习组件中使用套索回归对估计量方差与稳健性的影响。
- 为在现实世界因果推断中实施交叉拟合与平均提供实用指导。
提出的方法
- 本研究通过12种不同估计器的全面蒙特卡洛模拟进行,这些估计器在分割策略(两重、三重、五重)、交叉拟合与平均方法(均值、中位数、无平均)上存在差异。
- 每种估计器均独立应用于四种元学习器:T-learner、DR-learner、R-learner与X-learner,使用相同的数据生成过程(DGPs)。
- 对于交叉拟合,数据被划分为若干折, nuisance 函数在除用于CATE估计的那一折外的所有其他折上训练,结果在各折之间进行平均。
- 执行多次样本分割(最多50次),并取所得CATE估计的中位数以降低单次分割带来的方差与偏差。
- 通过均方误差(MSE)、平均绝对偏差与标准差等指标,在独立测试集上评估性能。
- 在部分变体中排除套索回归,以评估其对估计量稳定性与方差的影响。
实验结果
研究问题
- RQ1哪种样本分割与平均策略能在不同数据生成过程中最小化均方误差(MSE)?
- RQ2在机器学习组件中包含或排除套索回归如何影响估计量的方差与稳健性?
- RQ3与单次分割或均值平均方法相比,多次迭代中的中位数平均是否能提升估计量性能?
- RQ4在不同数据结构(随机对照试验与观察性研究)下,元学习器(T-learner、DR-learner、R-learner、X-learner)之间的性能差异如何变化?
- RQ5在有限样本中,为使中位数平均的性能稳定,最优迭代次数是多少?
主要发现
- 五重交叉拟合结合20次以上迭代的中位数平均,在所有数据生成过程与元学习器中均实现了最低的均方误差。
- 对于R-learner,在设置E中,中位数平均将MSE从5重交叉拟合的3.46降低至0.49,显示出显著的性能提升。
- 从机器学习组件中排除套索回归可显著降低方差,并在所有设定下实现更稳健且更具竞争力的结果。
- 在观察性研究中,DR-learner结合中位数平均优于其他估计器,而T-learner在某些设定下的MSE比DR-learner高出近2倍。
- X-learner在无分割的原始估计器下表现始终最佳,且中位数平均进一步降低了MSE,尤其在复杂、高依赖性设定下效果显著。
- 当存在异常值或重尾伪结果时,组合方法(交叉拟合 + 中位数平均)的方差可能高于仅使用交叉拟合,特别是在小样本或基于套索的模型中。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。