Skip to main content
QUICK REVIEW

[Paper Review] Interpretation of Linear Regression Coefficients under Mean Model Miss-Specification

Werner Brannath, Martin Scharpenberg|arXiv (Cornell University)|Sep 30, 2014
Advanced Statistical Methods and Models2 references3 citations
TL;DR

This paper proposes a new interpretation of linear regression coefficients under mean model miss-specification by defining a population-level measure of association that quantifies how much the expected outcome Y changes when the distribution of X is altered. The method uses a conservative linear approximation of this non-linear association measure, enabling robust, interpretable regression coefficients even when the true mean relationship is non-linear, with applications to confounding adjustment and inference under model misspecification.

ABSTRACT

Linear regression is a frequently used tool in statistics, however, its validity and interpretability relies on strong model assumptions. While robust estimates of the coefficients' covariance extend the validity of hypothesis tests and confidence intervals, a clear interpretation of the coefficients is lacking if the mean structure of the model is miss-specified. We therefore suggest a new intuitive and mathematical rigorous interpretation of the coefficients that is independent from specific model assumptions. It relies on a new population based measure of association. The idea is to quantify how much the population mean of the dependent variable Y can be changed by changing the distribution of the independent variable X. Restriction to linear functions for the distributional changes in X provides the link to linear regression. It leads to a conservative approximation of the newly defined and generally non-linear measure of association. The conservative linear approximation can then be estimated by linear regression. We show how this interpretation can be extended to multiple regression and how far and in which sense it leads to an adjustment for confounding. We point to perspectives for new analysis strategies and illustrate the utility and limitations of the new interpretation and strategies by examples and simulations.

Motivation & Objective

  • To address the lack of clear interpretation for linear regression coefficients when the mean model is miss-specified.
  • To develop a mathematically rigorous, population-based measure of association that is independent of specific model assumptions.
  • To provide a conservative linear approximation of this association measure using distributional changes in X.
  • To extend the interpretation to multiple regression and confounding adjustment in miss-specified models.
  • To enable valid inference and interpretation using standard linear regression even when the true relationship is non-linear.

Proposed method

  • Define a population measure of association ι_X(Y) as the maximal standardized covariance between Y and a distributional shift δ(X) of the covariate X.
  • Use the class of linear shifts δ(X) = a(X - E[X]) to obtain a conservative linear approximation of the non-linear association measure.
  • Show that the resulting coefficient corresponds to the least squares regression coefficient under weak regularity assumptions.
  • Derive the population parameter θ as the minimizer of the expected squared loss E[(Y - Xᵀθ)²], ensuring consistency under model misspecification.
  • Establish that the error term Ũ_i = Y_i - X_iᵀθ is uncorrelated with X_i, even when the true mean relationship is non-linear.
  • Extend the framework to multiple regression by conditioning on other covariates and defining partial association measures ι_{X_k|X_{-k}}(Y).

Experimental results

Research questions

  • RQ1What is a valid, interpretable interpretation of linear regression coefficients when the mean model is miss-specified?
  • RQ2How can a general, model-free measure of association between X and Y be defined using distributional shifts?
  • RQ3In what sense is the linear approximation of the non-linear association measure conservative and interpretable?
  • RQ4How can this framework be used to adjust for confounding in the presence of non-linear mean structures?
  • RQ5Under what conditions is the linear coefficient free of confounding bias in miss-specified models?

Key findings

  • The regression coefficient θ in miss-specified linear models consistently estimates the population parameter that minimizes the expected squared loss, even when the true mean function is non-linear.
  • The error term Ũ_i = Y_i - X_iᵀθ is uncorrelated with X_i, ensuring valid inference via Huber-White sandwich standard errors.
  • The proposed association measure ι_X(Y) quantifies the maximal change in E[Y] per unit change in the distribution of X, providing a causal-like interpretation.
  • The linear approximation of ι_X(Y) via δ(X) = a(X - E[X]) yields the ordinary least squares coefficient, justifying its use under model misspecification.
  • The partial association measure ι_{X_k|X_{-k}}(Y) is free of confounding if and only if E[X_k|X_{-k}] is linear in the other covariates.
  • Simulations and examples show that the new interpretation enables valid inference and confounding adjustment even when the true relationship is non-linear.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.