Skip to main content
QUICK REVIEW

[论文解读] Models as Approximations II: A Model-Free Theory of Parametric Regression

Andreas Buja, Lawrence Brown|arXiv (Cornell University)|Dec 10, 2016
Statistical Methods and Inference参考文献 25被引用 6
一句话总结

本文通过将回归参数重新定义为统计泛函(即从任意联合分布中导出的量,而非假设存在正确模型),发展了一套无模型的参数回归理论。该理论引入了一种基于回归变量分布不变性的新型‘良好设定’概念,通过重新加权实现诊断,并表明抽样变异性可分解为两个N⁻¹ᐟ²分量:一个来自条件响应分布,另一个来自回归变量分布与设定错误的交互作用。该框架支持使用x-y自助法或沙门德 estimator 进行稳健推断,证据表明自助法更具稳定性。

ABSTRACT

We develop a model-free theory of general types of parametric regression for iid observations. The theory replaces the parameters of parametric models with statistical functionals, to be called "regression functionals'', defined on large non-parametric classes of joint $\xy$ distributions, without assuming a correct model. Parametric models are reduced to heuristics to suggest plausible objective functions. An example of a regression functional is the vector of slopes of linear equations fitted by OLS to largely arbitrary $\xy$ distributions, without assuming a linear model (see Part~I). More generally, regression functionals can be defined by minimizing objective functions or solving estimating equations at joint $\xy$ distributions. In this framework it is possible to achieve the following: (1)~define a notion of well-specification for regression functionals that replaces the notion of correct specification of models, (2)~propose a well-specification diagnostic for regression functionals based on reweighting distributions and data, (3)~decompose sampling variability of regression functionals into two sources, one due to the conditional response distribution and another due to the regressor distribution interacting with misspecification, both of order $N^{-1/2}$, (4)~exhibit plug-in/sandwich estimators of standard error as limit cases of $\xy$ bootstrap estimators, and (5)~provide theoretical heuristics to indicate that $\xy$ bootstrap standard errors may generally be more stable than sandwich estimators.

研究动机与目标

  • 用一种不依赖模型假设的新概念——'良好设定'——取代传统回归泛函中对模型正确性的观念。
  • 开发一种诊断工具,通过重新加权回归变量分布来检测设定错误,适用于所有类型的变量,包括非定量变量。
  • 将回归泛函的抽样变异性分解为两个独立来源:条件响应分布和回归变量分布与设定错误的交互作用,两者均以N⁻¹ᐟ²速率收敛。
  • 通过证明x-y自助法标准误是插补/沙门德 estimator 的极限情况,建立理论基础,使稳健推断成为可能,并表明其可能更具稳定性。
  • 提供一个通用的参数回归框架,将模型视为近似而非真理,并通过参数的泛函解释支持局部建模。

提出的方法

  • 将回归泛函定义为从任意联合(x-y)分布到数值的映射,通过最小化目标函数或求解估计方程导出。
  • 将'良好设定'定义为:当回归变量的边缘分布发生变化时,回归泛函保持不变,使其成为仅依赖于条件响应分布的性质。
  • 提出一种基于重新加权的诊断方法,通过扰动回归变量分布(而不发生位移)来检测泛函对回归变量分布的依赖性,从而识别设定错误。
  • 将回归泛函的渐近抽样变异性分解为两个N⁻¹ᐟ²分量:一个源于条件响应分布,另一个源于回归变量分布与模型设定错误的交互作用。
  • 推导出插补/沙门德 estimator 作为x-y自助法估计量的极限情况,建立自助法与经典推断之间的理论联系。
  • 理论上论证,由于x-y自助法标准误具有非渐近性和基于重采样的本质,其一般而言可能比沙门德 estimator 更具稳定性。

实验结果

研究问题

  • RQ1在无模型框架下,如何在不假设存在真实模型的前提下,重新定义参数回归中的模型正确性概念?
  • RQ2当底层模型不正确时,回归泛函的'良好设定'意味着什么?
  • RQ3在不依赖参数假设或分布位移的前提下,如何诊断回归泛函的设定错误?
  • RQ4回归泛函中的抽样变异性在何种方式上可分解为与响应变量和回归变量分布相关的分量?
  • RQ5x-y自助法标准误与插补/沙门德 estimator 之间存在何种理论关系?在有限样本中,哪一种更具稳健性?

主要发现

  • 回归泛函的'良好设定'被定义为:当回归变量的边缘分布发生变化时,其保持不变,因此仅是条件响应分布的性质。
  • 回归泛函的抽样变异性可分解为两个N⁻¹ᐟ²分量:一个源于条件响应分布,另一个源于回归变量分布与设定错误的交互作用。
  • 重新加权诊断方法可通过评估泛函对回归变量分布的依赖性来检测设定错误,且该方法适用于所有变量类型,包括定性变量。
  • 已证明插补/沙门德 estimator 是x-y自助法估计量的极限情况,建立了重采样方法与经典推断之间的理论联系。
  • 理论启发表明,由于x-y自助法标准误具有非渐近性和数据驱动的特性,其一般而言可能比沙门德 estimator 更具稳定性。
  • 该框架可通过探索局部最优近似来实现模型的局部化,极大地扩展了参数模型在单一拟合之外的表达能力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。