[Paper Review] Shapley value confidence intervals for variable selection in regression models
This paper proposes asymptotic confidence intervals for Shapley values in regression models to assess variable importance, leveraging the coefficient of determination's decomposition via game theory. Under elliptical symmetry, it establishes asymptotic normality of Shapley values, enabling efficient hypothesis testing and interval estimation that outperform bootstrap methods in computation and accuracy.
Multiple linear regression is a commonly used inferential and predictive process, whereby a single response variable is modeled via an affine combination of multiple explanatory covariates. The coefficient of determination is often used to measure the explanatory power of the chosen combination of covariates. A ranking of the explanatory contribution of each of the individual covariates is often sought in order to draw inference regarding the importance of each covariate with respect to the response phenomenon. A recent method for ascertaining such a ranking is via the game theoretic Shapley value decomposition of the coefficient of determination. Such a decomposition has the desirable efficiency, monotonicity, and equal treatment properties. Under an elliptical assumption, we obtain the asymptotic normality of the Shapley values. We then utilize this result in order to construct confidence intervals and hypothesis tests regarding such quantities. Monte Carlo studies regarding our results are provided. We found that our asymptotic confidence intervals are computationally superior to competing bootstrap methods and are able to improve upon the performance of such intervals. Analyses of housing and real estate data are used to demonstrate the applicability of our methodology.
Motivation & Objective
- To address the need for statistical inference on variable importance in regression models beyond mere ranking.
- To provide a computationally efficient alternative to bootstrap-based confidence intervals for Shapley values.
- To establish asymptotic normality of Shapley values under elliptical distributional assumptions for valid inference.
- To enable hypothesis testing and confidence interval construction for individual variable contributions to R².
- To demonstrate the method’s practical utility through real-world housing and real estate data applications.
Proposed method
- The paper uses the game-theoretic Shapley value decomposition of the coefficient of determination (R²) to assign fair, efficient, and monotonic contributions to each covariate.
- It derives the asymptotic distribution of Shapley values under the assumption of elliptical symmetry in the covariates.
- Asymptotic normality of the Shapley values is established, enabling the construction of confidence intervals and hypothesis tests.
- The method relies on the delta method and variance-covariance estimation under regularity conditions for the asymptotic distribution.
- Monte Carlo simulations are used to validate the finite-sample performance of the confidence intervals.
- The approach is compared empirically to bootstrap-based intervals in terms of coverage probability and computational cost.
Experimental results
Research questions
- RQ1Can asymptotic confidence intervals for Shapley values be constructed under elliptical symmetry assumptions in regression models?
- RQ2How do these asymptotic intervals compare to bootstrap-based intervals in terms of coverage accuracy and computational efficiency?
- RQ3Do the proposed intervals maintain appropriate type I error rates in hypothesis testing for variable importance?
- RQ4What is the empirical performance of the method in real-world regression applications with real estate and housing data?
- RQ5Can the Shapley value decomposition of R² be used to provide valid statistical inference on variable contributions?
Key findings
- The Shapley values of regression coefficients are asymptotically normal under elliptical symmetry, enabling valid inference.
- The proposed asymptotic confidence intervals achieve comparable or better coverage accuracy than bootstrap methods.
- The asymptotic approach is significantly faster than bootstrap-based methods, especially in high-dimensional settings.
- Monte Carlo studies confirm that the asymptotic intervals maintain appropriate coverage levels across various sample sizes and correlation structures.
- In real data analyses, the method successfully identifies key predictors in housing and real estate datasets with reliable inference.
- The method improves upon existing approaches by combining theoretical rigor with computational efficiency for variable importance assessment.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.