Skip to main content
QUICK REVIEW

[论文解读] Estimating Mutual Information for Discrete-Continuous Mixtures

Weihao Gao, Sreeram Kannan|arXiv (Cornell University)|Sep 19, 2017
Bayesian Modeling and Causal Inference参考文献 5被引用 77
一句话总结

本文提出一种基于 Radon-Nikodym 导数和 k 最近邻的混合离散-连续分布的互信息估计量,证明一致性,并在基线方法上显示出有利性能。

ABSTRACT

Estimating mutual information from observed samples is a basic primitive, useful in several machine learning tasks including correlation mining, information bottleneck clustering, learning a Chow-Liu tree, and conditional independence testing in (causal) graphical models. While mutual information is a well-defined quantity in general probability spaces, existing estimators can only handle two special cases of purely discrete or purely continuous pairs of random variables. The main challenge is that these methods first estimate the (differential) entropies of X, Y and the pair (X;Y) and add them up with appropriate signs to get an estimate of the mutual information. These 3H-estimators cannot be applied in general mixture spaces, where entropy is not well-defined. In this paper, we design a novel estimator for mutual information of discrete-continuous mixtures. We prove that the proposed estimator is consistent. We provide numerical experiments suggesting superiority of the proposed estimator compared to other heuristics of adding small continuous noise to all the samples and applying standard estimators tailored for purely continuous variables, and quantizing the samples and applying standard estimators tailored for purely discrete variables. This significantly widens the applicability of mutual information estimation in real-world applications, where some variables are discrete, some continuous, and others are a mixture between continuous and discrete components.

研究动机与目标

  • 在变量为混合离散与连续时,推动准确的互信息估计。
  • 开发一种直接估计器,适用于 entropy 不易定义的一般测度空间。
  • 为估计量建立理论保证(一致性)。
  • 在合成数据和真实数据上对比标准基线,展示实际性能。

提出的方法

  • 通过 Radon-Nikodym 导数定义一般分布的 MI。
  • 提出一种混合变量 MI 估计量,利用 k-NN 距离在每个样本处估计导数。
  • 在一个统一方案中处理离散点、联合密度区域和纯连续部分。
  • 在温和的技术条件下证明估计量的一致性(ℓ2-一致性)。
  • 显示该估计量在极限情况下回收特殊情况(纯离散、纯连续或混合)。
  • 在实验中与离散化、自适应分区和噪声-KSG 基线进行对比。

实验结果

研究问题

  • RQ1当 X 和/或 Y 是离散和连续成分的混合时,MI 能否得到一致估计?
  • RQ2基于 Radon-Nikodym 的估计量在混合情形下相对于 3H 基于和 KSG 风格方法的性能如何?
  • RQ3所提出的估计量是否能够自适应纯离散、纯连续和混合设置?
  • RQ4在合成和真实的混合数据任务中,该估计量的经验性能如何?

主要发现

  • 在 ℓ2 下,在技术假设和常见实际分布下,估计量是一致的。
  • 在一系列实验中,它优于对混合变量进行离散化或添加噪声的基线方法。
  • 该方法自然地将纯离散和纯连续情况作为特殊实例回收。
  • 实验包括更高维的混合和零膨胀分布,显示出鲁棒性。
  • 应用包括在 dropout 异常污染下的特征选择和基因调控网络推断。
  • 经验结果显示相对于离线化和噪声-KSG 方法,样本效率更优。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。