Skip to main content
QUICK REVIEW

[Paper Review] Phase Transitions in Unsupervised Feature Selection

Jonathan Fiorentino, Michele Monti|arXiv (Cornell University)|Jan 31, 2026
Machine Learning in Bioinformatics0 citations
TL;DR

The paper analyzes an unsupervised feature selection pipeline based on Differentiable Information Imbalance (DII) applied to protein feature sets, revealing a phase-transition-like behavior that depends on feature type and correlational structure, and links critical feature counts to supervised classification performance.

ABSTRACT

Identifying minimal and informative feature sets is a central challenge in data analysis, particularly when few data points are available. Here we present a theoretical analysis of an unsupervised feature selection pipeline based on the Differentiable Information Imbalance (DII). We consider the specific case of structural and physico-chemical features describing a set of proteins. We show that if one considers the features as coordinates of a (hypothetical) statistical physics model, this model undergoes a phase transition as a function of the number of retained features. For physico-chemical descriptors, this transition is between a glass-like phase when the features are few and a liquid-like phase. The glass-like phase exhibits bimodal order-parameter distributions and Binder cumulant minima. In contrast, for structural descriptors the transition is less sharp. Remarkably, for physico-chemical descriptors the critical number of features identified from the DII coincides with the saturation of downstream binary classification performance. These results provide a principled, unsupervised criterion for minimal feature sets in protein classification and reveal distinct mechanisms of criticality across different feature types.

Motivation & Objective

  • Motivate unsupervised feature selection when labeled data are scarce.
  • Study how DII acts as an order parameter in selecting informative feature subsets.
  • Characterize how feature-set structure (physico-chemical vs structural) affects the information landscape.
  • Relate the unsupervised critical feature count to downstream binary classification performance.

Proposed method

  • Define and compute DII as an unsupervised order parameter for feature subsets.
  • Apply backward feature elimination using DII on physico-chemical and structural feature sets.
  • Analyze the distribution of DII values across random subsamples to study landscape ruggedness.
  • Use Binder cumulant analysis to identify a critical feature number indicating transition points.
  • Train a classifier (MLP) to relate feature count to binary classification performance AUROC.
Figure 1: Criticality in the Differentiable Information Imbalance during feature elimination. (A,B) Average DII versus the number of non-zero features $F$ for the LLPS dataset, for physico-chemical (A) and structural (B) features. (C,D) Heatmaps of the log-transformed probability density of the DII
Figure 1: Criticality in the Differentiable Information Imbalance during feature elimination. (A,B) Average DII versus the number of non-zero features $F$ for the LLPS dataset, for physico-chemical (A) and structural (B) features. (C,D) Heatmaps of the log-transformed probability density of the DII

Experimental results

Research questions

  • RQ1Does DII exhibit a phase-transition-like behavior as the number of retained features increases?
  • RQ2How does the nature of the feature set (physico-chemical vs structural) influence the type of transition (glass-like vs crossover)?
  • RQ3Is there a link between the unsupervised critical feature count and the saturation point of downstream classification performance?
  • RQ4How do correlations and variance heterogeneity in feature sets drive the information landscape?

Key findings

  • Physico-chemical features show a glass-like transition with bimodal DII landscapes and a Binder cumulant minimum.
  • Structural features display a weaker, smoother transition or crossover, with unimodal DII distributions.
  • Correlation structure drives the transition for physico-chemical features, while variance heterogeneity drives it for structural features.
  • The critical feature count for physico-chemical descriptors coincides with the saturation point of binary classification performance when using DII-selected features.
  • High-level: informative features behave as interacting degrees of freedom under constraints, connecting criticality to generalization in protein classification.
Figure 2: Binder cumulant analysis reveals a glass-like phase transition for physico-chemical features. (A,C) Binder cumulant $U(F)$ as a function of the number of non-zero features $F$ for physico-chemical (A) and structural (C) descriptors, for the LLPS dataset. (B,D) Extrapolation of $F_{min}$ (p
Figure 2: Binder cumulant analysis reveals a glass-like phase transition for physico-chemical features. (A,C) Binder cumulant $U(F)$ as a function of the number of non-zero features $F$ for physico-chemical (A) and structural (C) descriptors, for the LLPS dataset. (B,D) Extrapolation of $F_{min}$ (p

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.