[论文解读] Quantifying multivariate redundancy with maximum entropy decompositions of mutual information
本文提出了一种基于最大熵的框架,用于量化信息分解中的多变量冗余,通过有根树结构的约束将互信息分解为非负且符合公理的分量。该方法将双变量冗余度量推广至多变量设置,确保与冗余格结构一致,并能精确分离唯一性、冗余性和协同性信息贡献。
Williams and Beer (2010) proposed a nonnegative mutual information decomposition, based on the construction of redundancy lattices, which allows separating the information that a set of variables contains about a target variable into nonnegative components interpretable as the unique information of some variables not provided by others as well as redundant and synergistic components. However, the definition of multivariate measures of redundancy that comply with nonnegativity and conform to certain axioms that capture conceptually desirable properties of redundancy has proven to be elusive. We here present a procedure to determine nonnegative multivariate redundancy measures, within the maximum entropy framework. In particular, we generalize existing bivariate maximum entropy measures of redundancy and unique information, defining measures of the redundant information that a group of variables has about a target, and of the unique redundant information that a group of variables has about a target that is not redundant with information from another group. The two key ingredients for this approach are: First, the identification of a type of constraints on entropy maximization that allows isolating components of redundancy and unique redundancy by mirroring them to synergy components. Second, the construction of rooted tree-based decompositions of the mutual information, which conform to the axioms of the redundancy lattice by the local implementation at each tree node of binary unfoldings of the information using hierarchically related maximum entropy constraints. Altogether, the proposed measures quantify the different multivariate redundancy contributions of a nonnegative mutual information decomposition consistent with the redundancy lattice.
研究动机与目标
- 为解决长期存在的在信息分解中定义非负、符合公理的多变量冗余度量的挑战。
- 将双变量最大熵冗余度量扩展至多变量系统,同时保持与冗余格框架的一致性。
- 开发一种通过在分层结构约束下进行熵最大化来分离冗余和唯一冗余分量的方法。
- 确保所得分解满足冗余的关键公理,例如恒等公理,即使在存在确定性依赖关系的情况下也成立。
- 提供一种通用且可扩展的程序,通过有根树结构展开信息分量,实现多变量系统中互信息的分解。
提出的方法
- 利用特定条件独立性和互信息约束下的最大熵分布,以分离冗余和唯一冗余分量。
- 采用有根树结构的分解方法,其中每个节点通过分层相关的最大熵约束执行信息的二元展开。
- 通过将约束扩展至多变量设置,推广双变量冗余度量,确保非负性和公理合规性。
- 应用数学归纳法证明:当家族中至少存在一个分布对某组源变量的子集产生零信息时,实际分解与最大熵分解相等。
- 通过将随机分量与确定性分量分离,处理目标-源变量之间的确定性依赖关系,对随机部分应用相同的度量方法。
- 依赖互信息约束(例如,$C(X;i;j|k) = 0$)来定义冗余,针对因确定性重叠导致此类约束不可行的情况进行调整。
实验结果
研究问题
- RQ1如何以满足非负性和冗余格公理的方式量化多变量冗余?
- RQ2最大熵方法能否推广至多变量系统,以将互信息分解为非负且可解释的分量?
- RQ3对熵最大化的哪些约束允许分离冗余和唯一冗余分量,同时与协同分量保持对应关系?
- RQ4目标与源变量之间的确定性依赖关系如何影响冗余度量中互信息约束的有效性?
- RQ5该框架在多变量系统中可多大程度上通过信息分量的递归树状分解进行扩展?
主要发现
- 所提出的最大熵框架成功生成了符合冗余格结构的非负、符合公理的多变量冗余度量。
- 该方法通过有根树分解中的分层约束,将双变量冗余度量推广至多变量设置。
- 在家族中至少存在一个分布对某组源变量子集产生零信息的条件下,证明了实际分解与最大熵分解的相等性。
- 该框架确保冗余和唯一冗余项为非负,并与恒等公理一致,即使源变量在目标中部分或完全包含也成立。
- 对于具有确定性目标-源依赖关系的系统,当方法应用于冗余的随机分量时,其有效性依然成立,而确定性部分则单独处理。
- 该方法提供了一种可扩展、一致且基于信息论的多变量系统互信息分解方法,可将信息贡献精确分离为唯一性、冗余性和协同性分量。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。