Skip to main content
QUICK REVIEW

[Paper Review] Multiple factor analysis of distributional data

Rosanna Verde, Antonio Irpino|arXiv (Cornell University)|Apr 19, 2018
Image and Signal Denoising Methods22 references3 citations
TL;DR

This paper proposes a Multiple Factor Analysis (MFA) framework for distributional data using quantile variables derived from distribution-valued variables. By leveraging the squared $L_2$ Wasserstein distance as a metric, the method decomposes variability into components related to location, scale, and shape, enabling interpretable dimension reduction and visualization of distributions on factorial planes.

ABSTRACT

In the framework of Symbolic Data Analysis (SDA), distribution-variables are a particular case of multi-valued variables: each unit is represented by a set of distributions (e.g. histograms, density functions or quantile functions), one for each variable. Factor analysis (FA) methods are primary exploratory tools for dimension reduction and visualization. In the present work, we use Multiple Factor Analysis (MFA) approach for the analysis of data described by distributional variables. Each distributional variable induces a set new numeric variable related to the quantiles of each distribution. We call these new variables as extit{quantile variables} and the set of quantile variables related to a distributional one is a block in the MFA approach. Thus, MFA is performed on juxtaposed tables of quantile variables. \\ We show that the criterion decomposed in the analysis is an approximation of the variability based on a suitable metrics between distributions: the squared $L_2$ Wasserstein distance. \\ Applications on simulated and real distributional data corroborate the method. The interpretation of the results on the factorial planes is performed by new interpretative tools that are related to the several characteristics of the distributions (location, scale and shape).

Motivation & Objective

  • To extend Multiple Factor Analysis (MFA) to handle distributional data, where each individual is represented by a distribution (e.g., histogram, density, or quantile function).
  • To address the lack of a principled metric-based approach in existing PCA methods for distributional data, which often ignore scale and shape variability.
  • To develop a method that explicitly models variability using the $L_2$ Wasserstein distance between distributions, ensuring geometric and statistical consistency.
  • To provide interpretable visualization tools that link factorial plane structures to distributional characteristics such as location, scale, and shape.
  • To introduce novel interpretative tools—like the Spanish-fan plot—for understanding relationships among quantile variables in the reduced space.

Proposed method

  • Transform each distributional variable into a set of quantile variables, forming a block in MFA, where each quantile corresponds to a numeric variable.
  • Apply MFA on the juxtaposed table of quantile variables, with each block corresponding to one distributional variable.
  • Use the squared $L_2$ Wasserstein distance as the underlying metric to define variability, ensuring that the trace of the covariance matrix of quantile variables approximates the distributional variance.
  • Decompose total inertia into components related to position (mean), scale (spread), and shape (kurtosis, skewness) through interpretation of factorial axes.
  • Introduce the Spanish-fan plot to visualize the structure of quantile variables projected on factorial planes, linking fan shape to distributional features.
  • Perform dimension reduction by extracting principal components that capture the dominant sources of variation in the distributional data.

Experimental results

Research questions

  • RQ1How can Multiple Factor Analysis be adapted to handle distributional data represented by probability distributions?
  • RQ2What is the relationship between the variance of a distributional variable and the covariance structure of its quantile variables under the $L_2$ Wasserstein metric?
  • RQ3To what extent do the first factorial axes in MFA capture variability due to location, scale, and shape of distributions?
  • RQ4How can the results of MFA on distributional data be interpreted in terms of distributional characteristics like mean, spread, and kurtosis?
  • RQ5Can novel visualization tools such as the Spanish-fan plot effectively convey the structure of distributional data in factorial planes?

Key findings

  • The trace of the covariance matrix of quantile variables approximates the $L_2$ Wasserstein-based variance of the original distributional variable, validating the method’s metric consistency.
  • The first factorial axis is predominantly driven by differences in location (mean), with minimal influence from scale and shape components.
  • Scale and shape variations contribute only marginally to the total inertia, as evidenced by short eigenvalues on higher dimensions.
  • The Spanish-fan plot effectively visualizes the structure of quantile variables, with fan shape reflecting the kurtosis and spread of distributions.
  • In the BLOOD dataset, Hemoglobin and Hematocrit show strong positive correlation and higher means in younger individuals, consistent with biological expectations.
  • Kurtosis values for distributions (e.g., M-20 and F-20) show that distributions at the top of the factorial plane are less variable and less peaked than those at the bottom, confirming shape-based separation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.