Skip to main content
QUICK REVIEW

[论文解读] Interpretation of Linear Regression Coefficients under Mean Model Miss-Specification

Werner Brannath, Martin Scharpenberg|arXiv (Cornell University)|Sep 30, 2014
Advanced Statistical Methods and Models参考文献 2被引用 3
一句话总结

本文通过定义一种基于总体的关联度量,提出了一种在均值模型误设条件下的线性回归系数的新解释,该度量量化了当X的分布发生变化时,Y的期望值的变化程度。该方法使用该非线性关联度量的保守线性近似,即使真实均值关系为非线性,也能获得稳健且可解释的回归系数,适用于混杂因素调整和模型误设下的有效推断。

ABSTRACT

Linear regression is a frequently used tool in statistics, however, its validity and interpretability relies on strong model assumptions. While robust estimates of the coefficients' covariance extend the validity of hypothesis tests and confidence intervals, a clear interpretation of the coefficients is lacking if the mean structure of the model is miss-specified. We therefore suggest a new intuitive and mathematical rigorous interpretation of the coefficients that is independent from specific model assumptions. It relies on a new population based measure of association. The idea is to quantify how much the population mean of the dependent variable Y can be changed by changing the distribution of the independent variable X. Restriction to linear functions for the distributional changes in X provides the link to linear regression. It leads to a conservative approximation of the newly defined and generally non-linear measure of association. The conservative linear approximation can then be estimated by linear regression. We show how this interpretation can be extended to multiple regression and how far and in which sense it leads to an adjustment for confounding. We point to perspectives for new analysis strategies and illustrate the utility and limitations of the new interpretation and strategies by examples and simulations.

研究动机与目标

  • 解决当均值模型误设时,线性回归系数缺乏清晰解释的问题。
  • 开发一种数学上严谨、基于总体的关联度量,独立于特定模型假设。
  • 通过X的分布变化,对这一关联度量进行保守线性近似。
  • 将该解释扩展至多重回归及误设模型中的混杂因素调整。
  • 即使真实关系为非线性,仍能通过标准线性回归实现有效推断与可解释性。

提出的方法

  • 将总体关联度量 ι_X(Y) 定义为Y与协变量X的分布位移δ(X)之间的最大标准化协方差。
  • 使用线性位移类δ(X) = a(X - E[X]),以获得对非线性关联度量的保守线性近似。
  • 证明所得系数在弱正则性假设下对应于最小二乘回归系数。
  • 通过最小化期望平方损失 E[(Y - Xᵀθ)²] 的总体参数θ,确保在模型误设下的一致性。
  • 证明误差项 Ũ_i = Y_i - X_iᵀθ 与X_i不相关,即使真实均值关系为非线性。
  • 通过在其他协变量上条件化,并定义偏关联度量 ι_{X_k|X_{-k}}(Y),将该框架扩展至多重回归。

实验结果

研究问题

  • RQ1当均值模型误设时,线性回归系数的合理且可解释的解释是什么?
  • RQ2如何利用分布位移定义X与Y之间的一般性、模型无关的关联度量?
  • RQ3该非线性关联度量的线性近似在何种意义上是保守且可解释的?
  • RQ4该框架如何用于在非线性均值结构下调整混杂因素?
  • RQ5在何种条件下,误设模型中的线性系数可避免混杂偏倚?

主要发现

  • 在误设线性模型中,回归系数θ一致地估计了最小化期望平方损失的总体参数,即使真实均值函数为非线性。
  • 误差项 Ũ_i = Y_i - X_iᵀθ 与X_i不相关,确保可通过Huber-White异方差稳健标准误实现有效推断。
  • 所提出的关联度量 ι_X(Y) 量化了X分布每单位变化时E[Y]的最大变化,提供了类似因果的解释。
  • 通过δ(X) = a(X - E[X])对ι_X(Y)进行线性近似,可得到普通最小二乘系数,从而在模型误设下合理化其使用。
  • 偏关联度量 ι_{X_k|X_{-k}}(Y) 仅当E[X_k|X_{-k}]关于其他协变量为线性时,才可避免混杂。
  • 模拟与实例表明,新解释即使在真实关系为非线性时,也能实现有效推断与混杂因素调整。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。