[论文解读] Models as Approximations: How Random Predictors and Model Violations Invalidate Classical Inference in Regression
本文重新詮釋了Halbert White的穩健推斷框架,指出當模型為近似模型且預測變數為隨機時,傳統回歸推斷會失效。研究確立了在模型設定錯誤下的漸近理論下所導出的『沙 SALAD』估計量(sandwich estimator)能提供有效的推斷,而傳統標準誤差則因非線性與異質性而失效,兩者之間的差異在量值上可能任意大。
We review and interpret the early insights of Halbert White who over thirty years ago inaugurated a form of statistical inference for regression models that is asymptotically correct even under misspecification, that is, under the assumption that models are approximations rather than generative truths. This form of inference, which is pervasive in econometrics, relies on the of standard error. Whereas linear models theory in statistics assumes models to be true and predictors to be fixed, White's theory permits models to be approximate and predictors to be random. Careful reading of his work shows that the deepest consequences for statistical inference arise from a synergy --- a conspiracy --- of nonlinearity and randomness of the predictors which invalidates the ancillarity argument that justifies conditioning on the predictors when they are random. Unlike the standard error of linear models theory, the sandwich estimator provides asymptotically correct inference in the presence of both nonlinearity and heteroskedasticity. An asymptotic comparison of the two types of standard error shows that discrepancies between them can be of arbitrary magnitude. If there exist discrepancies, standard errors from linear models theory are usually too liberal even though occasionally they can be too conservative as well. A valid alternative to the sandwich estimator is provided by the bootstrap; in fact, the sandwich estimator can be shown to be a limiting case of the pairs bootstrap. We conclude by giving meaning to regression slopes when the linear model is an approximation rather than a truth. --- In this review we limit ourselves to linear least squares regression, but many qualitative insights hold for most forms of regression.
研究动机与目标
- 在模型為近似而非真正資料產生過程的脈絡下,重新表達Halbert White關於回歸穩健推斷的基礎性工作。
- 識別當預測變數為隨機且模型為非線性時,傳統推斷的關鍵失敗原因,挑戰標準線性模型理論中所使用的『輔助性』假設。
- 證明在模型設定錯誤下,傳統線性模型理論所產生的標準誤差通常過於寬鬆,且與『沙 SALAD』估計量之間的差異可能無界。
- 顯示自助法(特別是配對自助法)可作為『沙 SALAD』估計量的合理替代方案,且後者可視為前者的極限定義。
- 釐清當線性模型僅為近似而非真實資料產生過程時,回歸斜率的解釋意義。
提出的方法
- 應用模型設定錯誤下的漸近理論,將模型視為近似而非絕對正確的理論。
- 分析隨機預測變數與非線性對條件化於預測變數有效性的影響,從而動搖傳統推斷中核心的『輔助性』論證。
- 推導並比較兩種標準誤差:傳統線性模型標準誤差與White的『沙 SALAD』估計量,強調其在模型設定錯誤下的分歧。
- 透過漸近分析顯示,兩者標準誤差之間的差異可因非線性程度與異質性強度而任意變大。
- 證明『沙 SALAD』估計量是配對自助法在極限下的形式,從而建立重採樣與穩健變異數估計之間的理論連結。
- 依賴數學推理與漸近等價性,以支持在傳統推斷失效之情境下使用『沙 SALAD』估計量。
实验结果
研究问题
- RQ1當模型為近似且預測變數為隨機時,為何傳統回歸推斷會失效?
- RQ2在模型設定錯誤下,傳統標準誤差與『沙 SALAD』估計量之間差異的成因為何?
- RQ3非線性與隨機預測變數的協同效應如何使傳統推斷中所使用的『輔助性』論證失效?
- RQ4在模型設定錯誤下,『沙 SALAD』估計量如何作為傳統標準誤差的合理替代?
- RQ5當線性模型並非真實資料產生過程時,應如何詮釋回歸係數?
主要发现
- 基於固定預測變數與正確模型的傳統推斷,在預測變數為隨機且模型為近似時會失效,特別是因『輔助性』論證的破壞。
- 傳統標準誤差與『沙 SALAD』估計量之間的差異可能任意大,表示在模型設定錯誤下,傳統標準誤差通常過於寬鬆。
- 『沙 SALAD』估計量在異質性與模型非線性下均能提供漸近正確的推斷,使其對傳統假設的違反具有穩健性。
- 配對自助法在極限下收斂至『沙 SALAD』估計量,從而為後者提供理論基礎,使其可視為重採樣推斷的大型樣本近似。
- 即使線性模型非真實資料產生過程,回歸斜率仍可解釋為條件期望函數的局部近似。
- 核心洞見在於,模型設定錯誤與隨機預測變數的結合會破壞標準推斷,因而必須採用如『沙 SALAD』估計量等穩健替代方案。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。