[论文解读] Shapley value confidence intervals for variable selection in regression models
本文提出了一类用于回归模型中Shapley值的渐近置信区间,以评估变量重要性,利用博弈论方法对决定系数进行分解。在椭球对称假设下,建立了Shapley值的渐近正态性,从而实现高效的假设检验与区间估计,其计算效率和精度优于自助法(bootstrap)。
Multiple linear regression is a commonly used inferential and predictive process, whereby a single response variable is modeled via an affine combination of multiple explanatory covariates. The coefficient of determination is often used to measure the explanatory power of the chosen combination of covariates. A ranking of the explanatory contribution of each of the individual covariates is often sought in order to draw inference regarding the importance of each covariate with respect to the response phenomenon. A recent method for ascertaining such a ranking is via the game theoretic Shapley value decomposition of the coefficient of determination. Such a decomposition has the desirable efficiency, monotonicity, and equal treatment properties. Under an elliptical assumption, we obtain the asymptotic normality of the Shapley values. We then utilize this result in order to construct confidence intervals and hypothesis tests regarding such quantities. Monte Carlo studies regarding our results are provided. We found that our asymptotic confidence intervals are computationally superior to competing bootstrap methods and are able to improve upon the performance of such intervals. Analyses of housing and real estate data are used to demonstrate the applicability of our methodology.
研究动机与目标
- 为解决回归模型中变量重要性统计推断的需求,超越单纯的排序评估。
- 提供一种计算高效的替代方案,用于构建基于自助法的Shapley值置信区间。
- 在椭球分布假设下,建立Shapley值的渐近正态性,以支持有效的统计推断。
- 实现对R²中单个变量贡献的假设检验与置信区间构造。
- 通过真实房地产与住房数据应用,展示该方法的实际效用。
提出的方法
- 本文采用博弈论框架下的Shapley值分解方法,对决定系数(R²)进行分解,为每个协变量分配公平、高效且单调的贡献。
- 在协变量服从椭球对称分布的假设下,推导Shapley值的渐近分布。
- 建立Shapley值的渐近正态性,从而支持置信区间的构建与假设检验。
- 该方法依赖于在正则条件下对渐近分布的delta方法与方差-协方差估计。
- 通过蒙特卡洛模拟验证置信区间的有限样本性能。
- 在覆盖概率与计算成本方面,将该方法与基于自助法的区间进行实证比较。
实验结果
研究问题
- RQ1在回归模型中,基于椭球对称假设,能否构建Shapley值的渐近置信区间?
- RQ2与基于自助法的区间相比,这些渐近区间在覆盖准确性和计算效率方面表现如何?
- RQ3所提出的区间在变量重要性的假设检验中是否能保持适当的第一类错误率?
- RQ4该方法在真实房地产与住房数据的回归应用中表现如何?
- RQ5R²的Shapley值分解能否用于对变量贡献进行有效的统计推断?
主要发现
- 在椭球对称假设下,回归系数的Shapley值渐近服从正态分布,支持有效的统计推断。
- 所提出的渐近置信区间在覆盖准确度上与自助法相当或更优。
- 渐近方法在计算速度上显著优于基于自助法的方法,尤其在高维设置下优势明显。
- 蒙特卡洛研究证实,渐近区间在不同样本量和相关结构下均能保持适当的覆盖水平。
- 在真实数据分析中,该方法成功识别出住房与房地产数据集中关键预测变量,并提供可靠的统计推断。
- 该方法通过理论严谨性与计算高效性的结合,优于现有变量重要性评估方法。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。