[论文解读] Estimating Heterogeneous Causal Effects of High-Dimensional Treatments: Application to Conjoint Analysis
本文提出了一种贝叶斯有限高斯混合模型,用于在高维处理设置(如具有大量属性组合的联合实验)中估计异质性因果效应。通过基于处理效应模式对单位进行聚类,并利用协变量建模聚类成员关系,该方法识别出具有不同处理效应模式的受访者子群;在一项关于移民的联合调查中应用该方法,揭示出一个对非欧洲移民存在偏见的子群体。
Estimation of heterogeneous treatment effects is an active area of research. Most of the existing methods, however, focus on estimating the conditional average treatment effects of a single, binary treatment given a set of pre-treatment covariates. In this paper, we propose a method to estimate the heterogeneous causal effects of high-dimensional treatments, which poses unique challenges in terms of estimation and interpretation. The proposed approach finds maximally heterogeneous groups and uses a Bayesian mixture of regularized logistic regressions to identify groups of units who exhibit similar patterns of treatment effects. By directly modeling group membership with covariates, the proposed methodology allows one to explore the unit characteristics that are associated with different patterns of treatment effects. Our motivating application is conjoint analysis, which is a popular type of survey experiment in social science and marketing research and is based on a high-dimensional factorial design. We apply the proposed methodology to the conjoint data, where survey respondents are asked to select one of two immigrant profiles with randomly selected attributes. We find that a group of respondents with a relatively high degree of prejudice appears to discriminate against immigrants from non-European countries like Iraq. An open-source software package is available for implementing the proposed methodology.
研究动机与目标
- 填补因果推断方法在高维处理设置中的空白,其中处理组合数量远超样本量。
- 克服由因子设计引发的复杂处理效应异质性在估计与解释上的挑战。
- 通过识别具有相似处理效应模式的受访者子群,对个体层面的异质性进行建模。
- 实现对哪些协变量可预测具有不同处理效应特征的聚类成员关系的解释。
- 提供一种可扩展的、开源的解决方案,用于分析具有高维处理变异的联合实验。
提出的方法
- 使用贝叶斯有限高斯混合模型对多个单位聚类中的处理效应进行估计,模型基于正则化逻辑回归。
- 将聚类成员关系建模为预处理协变量的函数,以识别与不同效应模式相关的特征。
- 采用两步估计程序:首先在训练子样本上拟合模型以估计初始聚类分配;其次在完整数据上重新拟合,施加约束以提高稳定性。
- 引入正则化(如套索型惩罚)以处理高维处理效应并防止过拟合。
- 通过排列最小化实现标签切换校正,以在不同数据子样本间稳定聚类识别。
- 通过开源的 FactorHet R 包实现该方法,获取地址为 https://www.github.com/mgoplerud/FactorHet。
实验结果
研究问题
- RQ1在高维联合实验中,哪些受访者子群表现出系统性不同的处理效应?
- RQ2政治意识形态或民族优越感等协变量在多大程度上可预测属于具有不同处理效应模式的聚类成员?
- RQ3在每个识别出的聚类中,特定属性(如原籍国、合法身份)的处理效应大小和方向如何?
- RQ4在有限样本中,该方法与标准方法相比,在偏差、覆盖区间和均方误差方面表现如何?
- RQ5聚类结果在不同数据子样本和模型设定下有多大的稳定性?
主要发现
- 具有更高偏见水平的受访者子群对来自伊拉克等非欧洲国家的移民表现出显著的负面处理效应。
- 分样本重拟合方法相比完整数据估计,显著降低了偏差并改善了频率覆盖,尤其在真实效应非零时表现更优。
- 在较大样本量下,分样本方法的中位覆盖区间接近名义上的 95% 水平,仅有少数异常值显示覆盖不足。
- 原籍国的平均边际效应(MM)估计显示,在高偏见聚类中对非欧洲移民存在一致的歧视。
- 在样本划分中观察到稳定性问题,但通过排列最小化实现的标签切换校正有效稳定了各次重复中的聚类识别。
- 三聚类模型揭示,一个聚类始终对无证移民施加惩罚,而另一个聚类则对非欧洲国籍者表现出强烈偏见。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。