[论文解读] Regression-based causal inference with factorial experiments: estimands, model specifications, and design-based properties
本文提出了一种基于设计的框架,用于因子实验中的回归因果推断,表明基于因子的回归所得到的稳健标准误可对一般因子效应提供渐近保守的推断。该研究阐明了回归系数的因果解释,通过灵活加权定义了广义因子效应,并从设计基础视角量化了饱和与非饱和模型之间的偏差-方差权衡。
Factorial designs are widely used due to their ability to accommodate multiple factors simultaneously. The factor-based regression with main effects and some interactions is the dominant strategy for downstream data analysis, delivering point estimators and standard errors via one single regression. Justification of these convenient estimators from the design-based perspective requires quantifying their sampling properties under the assignment mechanism conditioning on the potential outcomes. To this end, we derive the sampling properties of the factor-based regression estimators from both saturated and unsaturated models, and demonstrate the appropriateness of the robust standard errors for the Wald-type inference. We then quantify the bias-variance trade-off between the saturated and unsaturated models from the design-based perspective, and establish a novel design-based Gauss--Markov theorem that ensures the latter's gain in efficiency when the nuisance effects omitted indeed do not exist. As a byproduct of the process, we unify the definitions of factorial effects in various literatures and propose a location-shift strategy for their direct estimation from factor-based regressions. Our theory and simulation suggest using factor-based inference for general factorial effects, preferably with parsimonious specifications in accordance with the prior knowledge of zero nuisance effects.
研究动机与目标
- 澄清因子实验中基于因子的模型里回归系数的因果解释。
- 通过任意加权方案定义广义因子效应,以提升外部有效性,并统一不同学科中现有的定义。
- 建立在大样本Wald型推断中,基于回归分析的稳健标准误的设计基础性质。
- 从设计基础视角量化饱和与非饱和回归模型之间的偏差-方差权衡。
- 为在不假设潜在结果模型的前提下使用最小二乘法结合稳健协方差矩阵进行推断,提供理论基础。
提出的方法
- 提出一种位置平移策略,通过最小二乘法将回归系数映射到广义因子效应的估计量。
- 在完全随机化下,推导回归估计量的设计基础抽样分布,条件于潜在结果。
- 使用Eicker–Huber–White稳健协方差矩阵作为真实抽样协方差的渐近保守估计量。
- 引入对比矩阵 $ G $ 以定义广义因子效应,允许任意加权方案。
- 在常数处理效应假设下,分析非饱和回归相对于饱和模型的偏差与方差。
- 在条件1下应用渐近理论,确保处理比例的收敛性以及潜在结果偏差的有界性。
实验结果
研究问题
- RQ1在因子实验的基于因子的模型中,回归系数的因果解释是什么?
- RQ2在此情境下,如何为大样本Wald型推断合理化最小二乘法所得的稳健标准误?
- RQ3从设计基础视角来看,饱和与非饱和回归模型之间的偏差-方差权衡是什么?
- RQ4如何定义广义因子效应,以统一因果推断、实验设计和社会科学文献中的标准定义?
- RQ5在何种条件下,非饱和回归相较于饱和模型具有更优的有限样本表现?
主要发现
- 在完全随机化下,基于因子的回归所得到的稳健标准误对真实抽样协方差具有渐近保守性,从而支持其在Wald型推断中的使用。
- 在位置平移设计基础框架下,包含主效应和选定交互作用的基于因子的模型中的回归系数,估计的是广义因子效应。
- 在常数处理效应下,非饱和回归的抽样方差小于饱和模型,但当存在干扰效应时会引入不衰减的偏差。
- 饱和与非饱和模型之间的偏差-方差权衡被正式量化,当干扰效应非零时,饱和模型更为稳妥。
- 该理论无需对潜在结果做建模假设,仅依赖于设计机制和条件1以保证渐近有效性。
- 该框架可经最小修改推广至一般的 $ Q_1 \times \cdots \times Q_K $ 因子设计,同时保持回归输出与矩估计量之间的对应关系。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。