Skip to main content
QUICK REVIEW

[论文解读] What can be estimated? Identifiability, estimability, causal inference and ill-posed inverse problems

Oliver J. Maclaren, Ruanui Nicholson|arXiv (Cornell University)|Apr 4, 2019
Bayesian Modeling and Causal Inference参考文献 41被引用 16
一句话总结

本文通过证明可识别性虽能保证因果 estimand 的唯一性,但无法保证其稳定性或实际可估性,重新定义了因果推断中可识别性与可估性之间的边界。作者使用抽象统计形式化和范畴论,证明了可识别量可能缺乏稳定估计量——使因果推断成为病态逆问题——从而主张真正的可估性必须满足希尔伯特(Hadamard)的三个条件:存在性、唯一性(可识别性)和稳定性。

ABSTRACT

We consider basic conceptual questions concerning the relationship between statistical estimation and causal inference. Firstly, we show how to translate causal inference problems into an abstract statistical formalism without requiring any structure beyond an arbitrarily-indexed family of probability models. The formalism is simple but can incorporate a variety of causal modelling frameworks, including 'structural causal models', but also models expressed in terms of, e.g., differential equations. We focus primarily on the structural/graphical causal modelling literature, however. Secondly, we consider the extent to which causal and statistical concerns can be cleanly separated, examining the fundamental question: 'What can be estimated from data?'. We call this the problem of estimability. We approach this by analysing a standard formal definition of 'can be estimated' commonly adopted in the causal inference literature -- identifiability -- in our abstract statistical formalism. We use elementary category theory to show that identifiability implies the existence of a Fisher-consistent estimator, but also show that this estimator may be discontinuous, and thus unstable, in general. This difficulty arises because the causal inference problem is, in general, an ill-posed inverse problem. Inverse problems have three conditions which must be satisfied to be considered well-posed: existence, uniqueness, and stability of solutions. Here identifiability corresponds to the question of uniqueness; in contrast, we take estimability to mean satisfaction of all three conditions, i.e. well-posedness. Lack of stability implies that naive translation of a causally identifiable quantity into an achievable statistical estimation target may prove impossible. Our article is primarily expository and aimed at unifying ideas from multiple fields, though we provide new constructions and proofs.

研究动机与目标

  • 澄清因果推断中可识别性与可估性之间的概念性区别。
  • 证明仅具备可识别性并不能确保因果查询在实际中可稳定或可行地进行统计估计。
  • 将因果推断视为受希尔伯特三个条件(存在性、唯一性、稳定性)支配的病态逆问题。
  • 通过抽象统计形式化,统一因果推断、逆问题与统计学习理论中的概念。
  • 主张可估性(而不仅仅是可识别性)应作为判断因果查询能否从数据中获得有意义估计的标准。

提出的方法

  • 在使用任意指标族的概率模型的抽象统计框架内形式化因果推断。
  • 应用初等范畴论证明:可识别性蕴含存在费雪一致估计量。
  • 表明此类估计量可能具有不连续性,因此在数据的微小扰动下不稳定。
  • 使用病态逆问题的框架分析因果估计,聚焦于希尔伯特的三个条件。
  • 引入“可估性”概念,即满足希尔伯特三个条件(存在性、唯一性(可识别性)和稳定性)的性质。
  • 通过计量经济学和因果建模中的例子(如影响函数无界的倾向得分法平均处理效应估计)说明估计量的不稳定性。

实验结果

研究问题

  • RQ1仅凭可识别性是否能确保因果查询可从数据中实际估计?
  • RQ2估计量的稳定性在多大程度上取决于因果模型和数据生成过程的结构?
  • RQ3病态逆问题的原则在统计与计量经济学中的因果推断问题中如何适用?
  • RQ4由于不稳定性,因果推断与统计推断之间的可分性假设在何种程度上会失效?
  • RQ5必须施加何种条件,才能确保一个可识别的因果量在实践中也具有可估性?

主要发现

  • 可识别性保证存在费雪一致估计量,但不能保证其稳定性,而稳定性对于实际估计至关重要。
  • 本文表明,可识别的因果 estimand 可能对应于不连续的估计量,导致其在数值上不稳定,且对数据的微小扰动缺乏鲁棒性。
  • 可估性不等同于可识别性;它要求满足希尔伯特的全部三个条件:存在性、唯一性(可识别性)和稳定性。
  • 来自计量经济学的例子(如通过倾向得分进行平均处理效应估计)表明,即使识别成立,无界的影响力函数也会导致不稳定性。
  • 识别算法的条件数可能极大,表明对输入扰动高度敏感,如半马尔可夫模型所示。
  • 来自逆问题理论、稳健统计学和学习理论的理论结果支持这一主张:只有具有稳定解的适定问题在原则上才是真正可估的。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。