[论文解读] Multiple factor analysis of distributional data
本文提出了一种基于分布值变量导出的分位数变量的多重因子分析(MFA)框架,用于分布数据。通过利用平方 $L_2$ Wasserstein 距离作为度量,该方法将变异分解为与位置、尺度和形状相关的分量,实现了分布在因子平面上的可解释性降维与可视化。
In the framework of Symbolic Data Analysis (SDA), distribution-variables are a particular case of multi-valued variables: each unit is represented by a set of distributions (e.g. histograms, density functions or quantile functions), one for each variable. Factor analysis (FA) methods are primary exploratory tools for dimension reduction and visualization. In the present work, we use Multiple Factor Analysis (MFA) approach for the analysis of data described by distributional variables. Each distributional variable induces a set new numeric variable related to the quantiles of each distribution. We call these new variables as extit{quantile variables} and the set of quantile variables related to a distributional one is a block in the MFA approach. Thus, MFA is performed on juxtaposed tables of quantile variables. \\ We show that the criterion decomposed in the analysis is an approximation of the variability based on a suitable metrics between distributions: the squared $L_2$ Wasserstein distance. \\ Applications on simulated and real distributional data corroborate the method. The interpretation of the results on the factorial planes is performed by new interpretative tools that are related to the several characteristics of the distributions (location, scale and shape).
研究动机与目标
- 将多重因子分析(MFA)扩展至处理分布数据,其中每个个体由一个分布(如直方图、密度函数或分位数函数)表示。
- 解决现有主成分分析(PCA)方法在分布数据中缺乏基于度量的系统性方法的问题,这些方法通常忽略尺度和形状的变异。
- 开发一种显式利用 $L_2$ Wasserstein 距离衡量分布间差异的方法,确保几何与统计上的一致性。
- 提供可解释的可视化工具,将因子平面上的结构与分布特征(如位置、尺度和形状)关联起来。
- 引入新型解释性工具(如西班牙风车图),以理解降维空间中分位数变量之间的关系。
提出的方法
- 将每个分布变量转换为一组分位数变量,构成 MFA 中的一个分块,其中每个分位数对应一个数值变量。
- 在分位数变量的拼接表上应用 MFA,每个分块对应一个分布变量。
- 使用平方 $L_2$ Wasserstein 距离作为基础度量以定义变异,确保分位数变量的协方差矩阵的迹近似于分布的方差。
- 通过因子轴的解释,将总惯性分解为与位置(均值)、尺度(离散程度)和形状(峰度、偏度)相关的分量。
- 引入西班牙风车图,以可视化分位数变量在因子平面上的结构,其扇形形状与分布特征相关联。
- 通过提取能捕捉分布数据中主要变异来源的主成分,实现降维。
实验结果
研究问题
- RQ1如何将多重因子分析适配于由概率分布表示的分布数据?
- RQ2在 $L_2$ Wasserstein 度量下,分布变量的方差与其实证分位数变量的协方差结构之间存在何种关系?
- RQ3MFA 的前几个因子轴在多大程度上捕捉了分布的位置、尺度和形状的变异?
- RQ4如何从分布特征(如均值、离散程度和峰度)的角度解释 MFA 在分布数据上的结果?
- RQ5诸如西班牙风车图等新型可视化工具是否能有效传达因子平面上分布数据的结构?
主要发现
- 分位数变量协方差矩阵的迹近似于原始分布变量基于 $L_2$ Wasserstein 距离的方差,验证了该方法在度量上的一致性。
- 第一个因子轴主要由位置差异(均值)驱动,尺度和形状分量的影响可忽略不计。
- 尺度和形状的变异对总惯性贡献甚微,表现为高维空间中特征值较短。
- 西班牙风车图能有效可视化分位数变量的结构,其扇形形状反映了分布的峰度与离散程度。
- 在 BLOOD 数据集中,血红蛋白与红细胞压积在年轻个体中表现出强正相关且均值较高,符合生物学预期。
- 分布的峰度值(如 M-20 和 F-20)表明,因子平面上方的分布比下方的分布变异更小、更平坦,证实了基于形状的分离。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。