[論文レビュー] Models as Approximations: How Random Predictors and Model Violations Invalidate Classical Inference in Regression
この論文はハルバート・ホワイトのロバスト推論フレームワークを再解釈し、モデルが近似であり予測変数が確率的である場合、古典的回帰推論が失敗することを示している。非線形性と異分散性のため、従来の標準誤差が失敗する状況において、モデルの不適合性下での漸近的理論から導かれるサンドイッチ推定量が有効な推論を提供することを確立している。推定量間の差異は、大きさが任意に取り得る可能性がある。
We review and interpret the early insights of Halbert White who over thirty years ago inaugurated a form of statistical inference for regression models that is asymptotically correct even under misspecification, that is, under the assumption that models are approximations rather than generative truths. This form of inference, which is pervasive in econometrics, relies on the of standard error. Whereas linear models theory in statistics assumes models to be true and predictors to be fixed, White's theory permits models to be approximate and predictors to be random. Careful reading of his work shows that the deepest consequences for statistical inference arise from a synergy --- a conspiracy --- of nonlinearity and randomness of the predictors which invalidates the ancillarity argument that justifies conditioning on the predictors when they are random. Unlike the standard error of linear models theory, the sandwich estimator provides asymptotically correct inference in the presence of both nonlinearity and heteroskedasticity. An asymptotic comparison of the two types of standard error shows that discrepancies between them can be of arbitrary magnitude. If there exist discrepancies, standard errors from linear models theory are usually too liberal even though occasionally they can be too conservative as well. A valid alternative to the sandwich estimator is provided by the bootstrap; in fact, the sandwich estimator can be shown to be a limiting case of the pairs bootstrap. We conclude by giving meaning to regression slopes when the linear model is an approximation rather than a truth. --- In this review we limit ourselves to linear least squares regression, but many qualitative insights hold for most forms of regression.
研究の動機と目的
- モデルを真のデータ生成過程ではなく近似として扱う文脈において、ハルバート・ホワイトの回帰におけるロバスト推論に関する基礎的業績を再表現すること。
- 予測変数が確率的でモデルが非線形である場合に、古典的推論が失敗する根本的要因を特定し、標準線形モデル理論で用いられる「付加的性」仮定の妥当性を問い直すこと。
- モデル不適合性下では、古典的線形モデル理論からの標準誤差がしばしば過度に楽観的であることを示し、サンドイッチ推定量との差異が無限大に達する可能性があること。
- ブートストラップ、特にペアスブートストラップがサンドイッチ推定量の有効な代替手段であることを示し、後者が極限的状態として現れることを示すこと。
- 線形モデルが真のデータ生成過程でない場合、回帰係数をどのように解釈すべきかを明確にすること。
提案手法
- モデル不適合性下での回帰モデルに漸近的理論を適用し、モデルを真の真理ではなく近似として扱う。
- 確率的予測変数と非線形性が、予測変数の条件付き推論の有効性に与える影響を分析し、古典的推論の中心的役割を果たす「付加的性」仮定を根底から揺るがす。
- 古典的線形モデルの標準誤差とホワイトのサンドイッチ推定量の2種類の標準誤差を導出し、比較することで、不適合性下での乖離を強調する。
- 漸近的分析を用いて、2つの標準誤差の差異が非線形性や異分散性の程度に応じて任意に大きくなる可能性があることを示す。
- サンドイッチ推定量がペアスブートストラップの極限形として現れることを示し、リサンプリングとロバスト分散推定の間の理論的関係を確立する。
- 数学的推論と漸近的同等性に依拠し、古典的推論が破綻する状況においてサンドイッチ推定量の使用を正当化する。
実験結果
リサーチクエスチョン
- RQ1モデルが近似であり予測変数が確率的である場合、線形回帰における古典的推論がなぜ失敗するのか。
- RQ2モデル不適合性下で、古典的標準誤差とサンドイッチ推定量との差異が生じる原因は何か。
- RQ3非線形性と確率的予測変数の相乗作用が、従来の推論で用いられる「付加的性」仮定をどのように無効にするのか。
- RQ4モデル不適合性下で、サンドイッチ推定量が古典的標準誤差の有効な代替手段であるという意味は何か。
- RQ5線形モデルが真のデータ生成過程でない場合、回帰係数はどのように解釈すべきか。
主な発見
- 予測変数が確率的でモデルが近似である場合、固定予測変数と真のモデルに基づく古典的推論は、特に「付加的性」仮定の失敗により崩壊する。
- 古典的標準誤差とサンドイッチ推定量との差異は、任意に大きくなる可能性があり、モデル不適合性下では古典的誤差がしばしば過度に楽観的である。
- サンドイッチ推定量は、異分散性およびモデル非線形性の両方の下で漸近的に正しい推論を提供し、古典的仮定の違反に強くロバストである。
- ペアスブートストラップは極限においてサンドイッチ推定量に収束し、後者の理論的基盤を、リサンプリングに基づく推論の漸近的近似として確立する。
- 線形モデルが真のデータ生成過程でない場合でも、回帰係数は条件付き平均関数の局所的近似として解釈可能である。
- 核心的な洞察は、モデル不適合性と確率的予測変数の併存が標準的推論を無効にし、サンドイッチ推定量のようなロバスト代替手段の必要性を生じることである。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。