[Paper Review] Estimating mutual information in high dimensions via classification error
This paper proposes a novel high-dimensional estimator of mutual information (MI) based on the average Bayes error from k-class classification, leveraging asymptotic theory under stratified sampling and high-dimensional limits. The method achieves superior performance over existing estimators by overcoming the logarithmic ceiling on MI estimation tied to class count k, enabling accurate MI estimation even in moderate dimensions when assumptions hold.
Multivariate pattern analyses approaches in neuroimaging are fundamentally concerned with investigating the quantity and type of information processed by various regions of the human brain; typically, estimates of classification accuracy are used to quantify information. While a extensive and powerful library of methods can be applied to train and assess classifiers, it is not always clear how to use the resulting measures of classification performance to draw scientific conclusions: e.g. for the purpose of evaluating redundancy between brain regions. An additional confound for interpreting classification performance is the dependence of the error rate on the number and choice of distinct classes obtained for the classification task. In contrast, mutual information is a quantity defined independently of the experimental design, and has ideal properties for comparative analyses. Unfortunately, estimating the mutual information based on observations becomes statistically infeasible in high dimensions without some kind of assumption or prior. In this paper, we construct a novel classification-based estimator of mutual information based on high-dimensional asymptotics. We show that in a particular limiting regime, the mutual information is an invertible function of the expected $k$-class Bayes error. While the theory is based on a large-sample, high-dimensional limit, we demonstrate through simulations that our proposed estimator has superior performance to the alternatives in problems of moderate dimensionality.
Motivation & Objective
- To address the limitations of classification accuracy as a proxy for mutual information in high-dimensional brain imaging data.
- To develop a theoretically grounded, classification-based estimator of mutual information that is invariant to arbitrary class partitioning.
- To overcome the fundamental limitation of existing estimators that are capped by log(k), where k is the number of classes.
- To provide a practical, scalable method for comparing information content across brain regions, modalities, or studies using mutual information.
- To establish conditions under which classification error can reliably estimate mutual information in high-dimensional settings.
Proposed method
- The method derives a theoretical relationship between mutual information and the expected k-class Bayes error under high-dimensional asymptotics and stratified sampling.
- It assumes that class labels are i.i.d. draws from a continuous distribution over an infinite number of potential classes, enabling control over classification error.
- The estimator uses a Poisson sampling approximation to model the probability of label collisions, leading to a stable estimate of the average Bayes error.
- It inverts the asymptotic relationship between Bayes error and mutual information to produce a plug-in estimator of MI.
- The approach relies on consistent classification error estimation, assuming the classifier approximates the Bayes rule.
- It is designed to be robust to arbitrary class partitioning and scalable to high-dimensional response spaces.
Experimental results
Research questions
- RQ1Can classification error in a k-class task be used to estimate mutual information in high-dimensional settings?
- RQ2How does the relationship between Bayes error and mutual information behave under high-dimensional asymptotics and stratified sampling?
- RQ3Can the estimator overcome the log(k) upper bound that limits existing classification-based MI estimators?
- RQ4How does the proposed estimator perform relative to established MI estimators in moderate-dimensional problems?
- RQ5Under what conditions is the estimator robust to model misspecification or finite-sample deviations?
Key findings
- The proposed estimator, denoted $\hat{I}_{HD}$, significantly outperforms existing estimators $\hat{I}_{Fano}$ and $\hat{I}_{CM}$ in simulations, especially in moderate dimensions.
- The estimator overcomes the fundamental limitation of previous methods that are capped at $\log(k)$ mutual information, even when the true MI is higher.
- In the worst-case example with perfect classification under partitioning, the estimator correctly identifies finite MI ($I(X;Y) = \log(k)$) without overestimation.
- The method maintains stable performance under varying numbers of classes, with no systematic increase or decrease in estimated MI, indicating robustness to class count.
- The estimator is most effective under the assumptions of stratified sampling and high-dimensional response space, with diagnostic checks suggested for real-world applicability.
- Theoretical justification is provided via a Poisson sampling model that approximates the average Bayes error, enabling inversion to estimate MI.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.