[论文解读] Stability theory of game-theoretic group feature explanations for machine learning models
本文为机器学习中的博弈论群体解释器构建了一个函数分析框架,引入了条件与边际博弈值的稳定性分析。提出了一类新型合作解释器,其联盟结构统一了两种解释类型,降低了复杂度,并通过参数化分层树推广至嵌套划分,展示了相较于传统方法在稳定性与理论鲁棒性上的改进。
In this article, we study feature attributions of Machine Learning (ML) models originating from linear game values and coalitional values defined as operators on appropriate functional spaces. The main focus is on random games based on the conditional and marginal expectations. The first part of our work formulates a stability theory for these explanation operators by establishing certain bounds for both marginal and conditional explanations. The differences between the two games are then elucidated, such as showing that the marginal explanations can become discontinuous on some naturally-designed domains, while the conditional explanations remain stable. In the second part of our work, group explanation methodologies are devised based on game values with coalition structure, where the features are grouped based on dependencies. We show analytically that grouping features this way has a stabilizing effect on the marginal operator on both group and individual levels, and allows for the unification of marginal and conditional explanations. Our results are verified in a number of numerical experiments where an information-theoretic measure of dependence is used for grouping.
研究动机与目标
- 将机器学习中的群体特征解释器形式化为函数分析空间中的算子,确保数学严谨性与可解释性。
- 在基于数据的度量下分析条件与边际博弈值的稳定性,揭示边际解释中的不稳定性。
- 通过具有结构化划分的合作博弈值统一条件与边际解释,降低复杂度并提升鲁棒性。
- 利用参数化分层树将分组方法推广至嵌套划分,支持递归合作博弈值。
- 建立合作博弈值的两步表示法,将其分解为中间博弈与基础值,实现理论与计算优势。
提出的方法
- 将解释算子形式化为从模型函数到特征贡献的映射,定义在关于数据与模型分布的 $ L^2 $ 空间上。
- 基于 $ \mathbb{E}[f(X)\mid X_S] $ 与 $ \mathbb{E}[f(x_S,X_{-S})]\big|_{x_S=X_S} $ 分别引入条件与边际博弈,以建模不同类型的解释。
- 提出扩展博弈值 $ \bar{h} $ 与 $ \bar{g} $ 以处理非合作博弈,实现稳定且可解释的群体解释。
- 开发合作博弈值的两步表示法:$ h^{(1)} $、$ h^{(2)} $ 与中间博弈 $ \hat{v}_T $,支持递归分解。
- 引入参数化分层树 $ \mathcal{T} $ 以建模嵌套分组,基于分层联盟定义递归解释器 $ \bar{u}_{S_j^\alpha} $。
- 应用基于互信息的预测器分组方法构建数据驱动的划分,提升群体解释器的稳定性与可解释性。
实验结果
研究问题
- RQ1在自然的数据基度量下,条件与边际博弈值的稳定性特性有何差异?
- RQ2基于结构化划分的合作博弈值能否统一条件与边际解释,同时降低计算复杂度?
- RQ3基于分层树的递归群体解释器的理论基础是什么?
- RQ4通过互信息进行预测器分组如何影响群体特征解释的稳定性与准确性?
- RQ5合作博弈值的两步分解是什么?其如何推广至嵌套划分?
主要发现
- 边际博弈值算子 $ \bar{\mathcal{E}}^{\text{ME}} $ 在自然的 $ L^2(P_X) $-基度量下被证明不稳定,而条件博弈值算子 $ \bar{\mathcal{E}}^{\text{CE}} $ 保持有界且连续。
- 基于 $ \bar{g} $ 的所提合作解释器统一了条件与边际解释,在计算复杂度上低于独立方法。
- 合作博弈值的两步表示法将任一此类值分解为两个基础值与一族中间博弈,支持模块化构建与分析。
- 基于参数化分层树 $ \mathcal{T} $ 定义的递归合作解释器将分组方法推广至嵌套划分,支持分层解释。
- 基于互信息的预测器分组显著提升了所得解释算子的稳定性,尤其在高维设置下表现突出。
- 扩展博弈值 $ \bar{g} $ 确保了在不同划分下解释算子的有界性与连续性,验证了其理论鲁棒性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。