Skip to main content
QUICK REVIEW

[論文レビュー] Necessary and Probably Sufficient Test for Finding Valid Instrumental Variables

Amit Sharma|arXiv (Cornell University)|Dec 4, 2018
Economic Policies and Impacts参考文献 2被引用数 6
ひとこと要約

本稿では、有効なIVと無効なIVの因果モデルの周辺尤度を比較することで、IVの妥当性を評価するベイジアンモデル比較手法—必須かつおそらく十分(NPS)検定—を提案する。離散変数におけるデータ駆動型のIV妥当性検出を可能とし、除外違反の検出において高い検出力を持ち、単調性および中程度~弱い道具具強度下で最も優れた性能を示す。

ABSTRACT

Can instrumental variables be found from data? While instrumental variable (IV) methods are widely used to identify causal effect, testing their validity from observed data remains a challenge. This is because validity of an IV depends on two assumptions, exclusion and as-if-random, that are largely believed to be untestable from data. In this paper, we show that under certain conditions, testing for instrumental variables is possible. We build upon prior work on necessary tests to derive a test that characterizes the odds of being a valid instrument, thus yielding the name "necessary and probably sufficient". The test works by defining the class of invalid-IV and valid-IV causal models as Bayesian generative models and comparing their marginal likelihood based on observed data. When all variables are discrete, we also provide a method to efficiently compute these marginal likelihoods. We evaluate the test on an extensive set of simulations for binary data, inspired by an open problem for IV testing proposed in past work. We find that the test is most powerful when an instrument follows monotonicity---effect on treatment is either non-decreasing or non-increasing---and has moderate-to-weak strength; incidentally, such instruments are commonly used in observational studies. Among as-if-random and exclusion, it detects exclusion violations with higher power. Applying the test to IVs from two seminal studies on instrumental variables and five recent studies from the American Economic Review shows that many of the instruments may be flawed, at least when all variables are discretized. The proposed test opens the possibility of data-driven validation and search for instrumental variables.

研究の動機と目的

  • 観測データからの道具具変数(IV)の妥当性を検証するという長年の課題に取り組むこと。標準的な仮定(除外仮定および仮想的独立性)は通常検証不能である。
  • 観測データのみを用いて有効な道具具と無効な道具具を区別できる統計的検定を開発すること。検証不能な分野的仮定に依存しない。
  • 同じデータセット内の複数の道具具の相対的妥当性を比較・評価するための体系的かつ客観的な基準を提供すること。
  • 道具具変数が離散的であり、単調性が成り立つ場合に、IV妥当性検定が実行可能となる条件を調査すること。
  • 広範なシミュレーションと実世界の応用を通じて、検定の実証的性能を評価し、広く使われている道具具に潜在する欠陥を明らかにすること。

提案手法

  • 有効なIVおよび無効なIVの因果モデルを、離散確率変数上のベイジアン生成モデルとして形式化する。
  • ディリクレ積分を用いて、離散設定下での観測データの周辺尤度を正確に計算し、有効なIVおよび無効なIVのモデルクラスの両方で周辺尤度を算出する。
  • 仮想的独立性または両方の除外仮定と仮想的独立性が破綻している場合の周辺尤度の閉形式表現を導出。ガンマ関数およびディリクレ積分を活用する。
  • 有効なIVモデルと無効なIVモデルの周辺尤度の比である妥当性比を、道具具の妥当性を評価するための検定統計量として用いる。
  • 道具具強度および違反レベルのさまざまな設定下で、シミュレートされたバイナリデータにNPS検定を適用し、検出力およびロバストネスを評価する。
  • 構造的モデルにおける単調性制約を用いて、理論的不等式を導出し、検定の設計を支援するとともに検出力の向上を図る。

実験結果

リサーチクエスチョン

  • RQ1観測データから道具具変数の妥当性を検証できる条件は何か。特に、未観測の交絡要因が存在する場合でも。
  • RQ2ベイジアンモデル比較をどのように用いることで、有効なIVモデルと無効なIVモデルの相対的妥当性を客観的に評価できるか。
  • RQ3提案された検定は、特に単調性下で、除外違反および仮想的独立性違反を検出する際にどの程度の検出力を持つのか。
  • RQ4道具具強度および単調性は、NPS検定の性能にどのように影響を与えるか。
  • RQ5経済学研究で広く使われている道具具が、NPSフレームワークで検証された場合、どの程度の確率で依然として妥当性を保っているのか。

主な発見

  • NPS検定は、特に単調性が成り立ち、道具具強度が中程度~弱い場合に、除外違反の検出において高い統計的検出力を達成する。
  • 単調性下で検定は最も優れた性能を示す。これは、道具具が処置に与える影響が、レベル間で非減少または非増加であることを意味する。
  • 除外仮定および仮想的独立性の両方が破綻している場合、ディリクレ積分による周辺尤度の閉形式計算のおかげで、NPS検定は強力な性能を維持する。
  • 周辺尤度に基づく妥当性比は、道具具の妥当性を比較するための原則的かつ解釈可能な指標を提供する。
  • アメリカン・エコノミック・レビューの古典的および最近の研究へのNPS検定の適用から、多くの広く使われている道具具が離散化された段階で無効である可能性が示唆された。
  • 本手法は、データ駆動型の道具具の妥当性検証および探索を可能とし、定性的な分野的根拠に依存する代替手段として体系的なアプローチを提供する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。