[Paper Review] Predictor-dependent shrinkage for linear regression via partial factor modeling
This paper proposes a partial factor model for linear regression in high-dimensional settings (p ≫ n) that improves predictor-dependent shrinkage by decoupling the conditional regression from the marginal predictor distribution. By reparameterizing the joint model hierarchically, it enables robust shrinkage through borrowing of information via hierarchical priors, outperforming standard methods like ridge regression and principal component regression across diverse covariance structures.
In prediction problems with more predictors than observations, it can sometimes be helpful to use a joint probability model, $π(Y,X)$, rather than a purely conditional model, $π(Y \mid X)$, where $Y$ is a scalar response variable and $X$ is a vector of predictors. This approach is motivated by the fact that in many situations the marginal predictor distribution $π(X)$ can provide useful information about the parameter values governing the conditional regression. However, under very mild misspecification, this marginal distribution can also lead conditional inferences astray. Here, we explore these ideas in the context of linear factor models, to understand how they play out in a familiar setting. The resulting Bayesian model performs well across a wide range of covariance structures, on real and simulated data.
Motivation & Objective
- To address the challenge of predictor-dependent shrinkage in high-dimensional linear regression (p ≫ n) where marginal predictor distributions can mislead conditional inference.
- To overcome the sensitivity of joint factor models to prior specification of the number of factors k, especially when k is misspecified or too small.
- To develop a method that borrows information from the predictor structure without letting the high-dimensional marginal distribution dominate posterior inference on regression coefficients.
- To provide a robust Bayesian framework that separates the estimation of the conditional regression from the marginal predictor model, ensuring reliable uncertainty quantification.
Proposed method
- Proposes a hierarchical partial factor model that reparameterizes the joint distribution π(Y, X) as π(X|θ_X)π(Y|X, θ_X, θ_Y), where θ_X governs the predictor structure and θ_Y governs the regression.
- Uses a compositional representation where the conditional mean E(Y|X) depends on both the latent factors f and the residual X − E(X|f), enabling flexible, data-driven shrinkage.
- Implements a hierarchical prior π(θ_Y|θ_X)π(θ_X) to allow information borrowing from the predictor structure without fixing the joint model structure.
- Introduces a prior on the rank of the loading matrix B_Y associated with the response, enabling estimation of the sufficient dimension of the regression subspace.
- Employs posterior computation via MCMC to estimate Pr(Λ = 0|X, Y), the probability that the response depends only on a subspace of the factors, thus identifying the effective dimension.
- Uses the posterior distribution of θ to perform shrinkage while maintaining uncertainty quantification, avoiding the pitfalls of fixed-k factor models.
Experimental results
Research questions
- RQ1How can we improve conditional regression inference in high-dimensional settings (p ≫ n) by leveraging the joint distribution of predictors and response?
- RQ2What are the consequences of misspecifying the number of latent factors k in a joint factor model, particularly when the response is correlated with low-importance components?
- RQ3Can we construct a joint model that allows for robust shrinkage in regression without being dominated by the high-dimensional marginal predictor distribution?
- RQ4How can we design a hierarchical prior structure that enables information borrowing from the predictor space while preserving flexibility in the conditional regression model?
- RQ5What is the role of the sufficient dimension reduction subspace in improving prediction and uncertainty quantification under model misspecification?
Key findings
- The partial factor model outperforms standard methods such as ridge regression, principal component regression, and least angle regression across both simulated and real data under diverse covariance structures.
- Posterior inference on the rank of the response-relevant subspace (via Pr(Λ = 0|X, Y)) effectively identifies the true effective dimension of the regression, even when k is misspecified.
- The method successfully mitigates the risk of overfitting to the marginal predictor distribution by decoupling the regression from the factor model's prior assumptions.
- The hierarchical prior structure allows for robust shrinkage without requiring precise specification of k, reducing sensitivity to prior choice.
- The approach maintains reliable uncertainty quantification in predictions by propagating uncertainty about the subspace dimension and factor structure through the posterior.
- The model’s flexibility is demonstrated by its ability to generalize to nonlinear settings through a conditional mean structure like E(Y|f,X) = g(f) + h(X − E(X|f)), suggesting broader applicability.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.