Skip to main content
QUICK REVIEW

[论文解读] Conditional partial exchangeability: a probabilistic framework for multi-view clustering

Beatrice Franzolini, Maria De Iorio|arXiv (Cornell University)|Jul 3, 2023
Bayesian Methods and Mixture ModelsComputer Science被引用 3
一句话总结

该论文提出条件部分可交换性(conditional partial exchangeability),这是一种新颖的贝叶斯非参数框架,用于多视图聚类,允许聚类结构在不同特征间变化,同时建模个体内部的依赖关系。通过利用层次化随机划分和依赖狄利克雷过程,该方法实现了灵活且计算可行的聚类,能够捕捉各特征的特异性贡献以及跨视图的异质聚类形状。

ABSTRACT

Standard clustering techniques assume a common configuration for all features in a dataset. However, when dealing with multi-view or longitudinal data, the clusters' number, frequencies, and shapes may need to vary across features to accurately capture dependence structures and heterogeneity. In this setting, classical model-based clustering fails to account for within-subject dependence across domains. We introduce conditional partial exchangeability, a novel probabilistic paradigm for dependent random partitions of the same objects across distinct domains. Additionally, we study a wide class of Bayesian clustering models based on conditional partial exchangeability, which allows for flexible dependent clustering of individuals across features, capturing the specific contribution of each feature and the within-subject dependence, while ensuring computational feasibility.

研究动机与目标

  • 解决标准聚类方法假设所有特征共享单一、统一聚类配置的局限性,该假设无法捕捉多视图或纵向数据中的异质性。
  • 在多个数据领域(如生长轨迹、代谢物、母亲临床数据)中建模个体内部依赖关系,其中各特征的聚类结构可能不同。
  • 构建一个概率框架,允许根据数据灵活调整各特征的聚类数量、形状和频率。
  • 通过构建具有条件部分可交换性的层次化贝叶斯模型,确保在高维、多视图设置下的计算可行性。
  • 通过使每个特征按比例贡献于聚类结构,克服标准聚类方法对高维特征的偏向。

提出的方法

  • 提出条件部分可交换性作为跨多个领域依赖随机划分的概率范式,允许各特征的聚类配置不同,同时保持个体间的依赖关系。
  • 构建一个层次化贝叶斯模型,其中第一层定义各特征的聚类配置,第二层通过共享潜在结构来建模特征间聚类分配的依赖关系。
  • 使用依赖狄利克雷过程先验来建模各特征间聚类分配的联合分布,确保各视图中聚类成员关系在统计上相互依赖。
  • 引入特征特异性似然函数(如多元正态分布),并通过预测分布对聚类配置进行积分,以计算边际似然。
  • 通过条件化第一层聚类配置并积分第二层依赖关系,推导出联合预测分布,从而实现完整的贝叶斯推断。
  • 通过利用条件独立结构并使用预测递归方法,高效计算边际似然,确保计算上的可处理性。
(a) True clustering structure. Observations in each cluster are simulated from a Multivariate Normal with identity variance and covariance matrix.
(a) True clustering structure. Observations in each cluster are simulated from a Multivariate Normal with identity variance and covariance matrix.

实验结果

研究问题

  • RQ1如何在保持特征间个体依赖关系的前提下,建模在多个特征或视图间变化的聚类结构?
  • RQ2何种概率框架可实现灵活的、特征特异性的聚类形状与数量,而无需假设统一的聚类配置?
  • RQ3如何确保在多视图数据中,高维特征不会主导聚类结果?
  • RQ4我们能否构建一个贝叶斯非参数模型,支持跨视图的个体内部依赖关系,同时允许随机、数据驱动的聚类选择?
  • RQ5条件部分可交换性在支持具有异质数据类型的多个领域间依赖聚类中起到何种作用?

主要发现

  • 所提出的框架在玩具示例中成功捕捉了各特征间不同的聚类结构,其中一维和三维特征表现出截然不同的聚类模式。
  • 标准聚类方法如k-means和狄利克雷过程混合模型在玩具示例中未能检测到真实聚类结构,因其被高维特征主导。
  • 该模型的预测分布考虑了特征特异性贡献和个体内部依赖关系,避免了对高维特征的偏向。
  • 条件部分可交换性实现了对各特征间聚类分配的联合建模,当可交换性成立时,联合概率在特定置换下保持不变。
  • 该框架通过非参数先验支持各特征聚类数量的自动选择,无需预先指定聚类数量。
  • 理论分析证实,该模型在特定配置下满足条件可交换性,确保在依赖视图间进行有效概率推断。
(b) k-means clustering configuration with the number of clusters determined by elbow plot, gap statistics ( Tibshirani et al. , 2001 ) , and silhouette method ( Kaufman and Rousseeuw , 2009 ) .
(b) k-means clustering configuration with the number of clusters determined by elbow plot, gap statistics ( Tibshirani et al. , 2001 ) , and silhouette method ( Kaufman and Rousseeuw , 2009 ) .

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。