Skip to main content
QUICK REVIEW

[论文解读] Two-sample instrumental variable analyses using heterogeneous samples

Qingyuan Zhao, Jingshu Wang|arXiv (Cornell University)|Aug 31, 2017
Advanced Causal Inference Techniques参考文献 53被引用 10
一句话总结

本文提出了一类在线性结构模型下对异质样本具有鲁棒性的新型两样本工具变量(TSIV)估计量,表明尽管两阶段最小二乘法(TSLS)估计量并非渐近有效,但在实践中其表现几乎与最优估计量相当。当工具变量分布在不同样本间存在差异时,只要满足结构不变性和噪声同质性,该方法可确保因果估计的一致性。

ABSTRACT

Instrumental variable analysis is a widely used method to estimate causal effects in the presence of unmeasured confounding. When the instruments, exposure and outcome are not measured in the same sample, Angrist and Krueger (1992) suggested to use two-sample instrumental variable (TSIV) estimators that use sample moments from an instrument-exposure sample and an instrument-outcome sample. However, this method is biased if the two samples are from heterogeneous populations so that the distributions of the instruments are different. In linear structural equation models, we derive a new class of TSIV estimators that are robust to heterogeneous samples under the key assumption that the structural relations in the two samples are the same. The widely used two-sample two-stage least squares estimator belongs to this class. It is generally not asymptotically efficient, although we find that it performs similarly to the optimal TSIV estimator in most practical situations. We then attempt to relax the linearity assumption. We find that, unlike one-sample analyses, the TSIV estimator is not robust to misspecified exposure model. Additionally, to nonparametrically identify the magnitude of the causal effect, the noise in the exposure must have the same distributions in the two samples. However, this assumption is in general untestable because the exposure is not observed in one sample. Nonetheless, we may still identify the sign of the causal effect in the absence of homogeneity of the noise.

研究动机与目标

  • 解决当样本来自不同总体且工具变量分布异质时,传统两样本IV(TSIV)估计量存在的偏差问题。
  • 基于广义矩法(GMM)开发一类新的TSIV估计量,确保在样本间结构不变性条件下保持一致性。
  • 研究在放松线性假设时,TSIV估计量对模型误设和异质性的稳健性。
  • 澄清两样本设定下因果效应识别的条件,特别是误差分布同质性这一难以检验的假设。
  • 利用UK Biobank数据中子样本群体的真实数据,评估TSIV估计量在真实世界遗传流行病学应用中的表现。

提出的方法

  • 在工具变量、暴露变量和结果变量之间建立线性结构方程模型,假设两个异质样本间结构不变。
  • 利用独立的工具变量-暴露样本和工具变量-结果样本的样本矩,推导基于GMM的TSIV估计量族。
  • 建立在样本异质性下两样本TSLS估计量保持一致的条件,依赖于暴露模型的结构不变性。
  • 在GMM框架内识别最优TSIV估计量,该估计量在模型设定正确时可实现渐近效率。
  • 将分析扩展至非线性设定,表明除非误差分布同质,否则TSIV估计量对暴露模型误设不具稳健性。
  • 通过模拟和真实数据(UK Biobank)比较在样本异质条件下TSLS与最优TSIV估计量的表现。

实验结果

研究问题

  • RQ1当两样本来自不同总体且工具变量分布异质时,两样本工具变量估计量是否仍能保持一致性?
  • RQ2在样本异质性条件下,两阶段最小二乘法(TSLS)估计量与最优TSIV估计量在效率和偏差方面表现如何比较?
  • RQ3在结构不变性之外,两样本IV分析中因果效应估计所需的关键识别假设是什么?
  • RQ4当线性假设被放宽时,TSIV估计量对暴露模型误设的稳健性如何?
  • RQ5在两样本IV设定中,暴露方程中误差分布同质性的假设是否可检验,其违反会产生何种后果?

主要发现

  • 尽管TSLS估计量并非渐近有效,但在所有数值示例中(包括UK Biobank的真实数据),其表现几乎与最优TSIV估计量完全一致。
  • 最优TSIV估计量通过GMM推导得出,在模型设定正确且样本间结构不变时可实现渐近效率。
  • 在样本异质性存在时,TSIV估计量仅在暴露模型的结构在样本间保持不变(假设4)时才保持一致。
  • 当暴露模型为非线性或设定错误时,TSIV估计量会产生偏差,凸显其与单样本IV分析相比的关键脆弱性。
  • 暴露方程中误差分布同质性的假设对于非线性识别是必要的,但由于一个样本中未观测到暴露变量,该假设无法检验。
  • 在使用UK Biobank的真实数据分析中,来自异质样本(如子样本群体)的TSIV估计值与基准值存在差异,但因标准误增大而未达显著水平。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。