Skip to main content
QUICK REVIEW

[论文解读] Community models for partially observed networks from surveys

Tianxi Li, Elizaveta Levina|arXiv (Cornell University)|Aug 9, 2020
Complex Network Analysis Techniques参考文献 34被引用 5
一句话总结

本文提出了一种针对通过调查采样获得的部分可观测网络的社区检测通用模型,通过个体偏好和社区结构建模边采样概率。该模型支持一致的谱聚类,并提出一种参数化的提名随机块模型用于有向网络,具备理论保证,且可通过矩方法实现高效估计。

ABSTRACT

Communities are a common and widely studied structure in networks, typically under the assumption that the network is fully and correctly observed. In practice, network data are often collected through sampling schemes such as surveys. These sampling mechanisms introduce noise and bias which can obscure the community structure and invalidate assumptions underlying standard community detection methods. We propose a general model for a class of network sampling mechanisms based on survey designs, designed to enable more accurate community detection for network data collected in this fashion. We model edge sampling probabilities as a function of both individual preferences and community parameters, and show community detection can be done by spectral clustering under this general class of models. We also propose, as a special case of the general framework, a parametric model for directed networks we call the nomination stochastic block model, which allows for meaningful parameter interpretations and can be fitted by the method of moments. Both spectral clustering and the method of moments in this case are computationally efficient and come with theoretical guarantees of consistency. We evaluate the proposed model in simulation studies on both unweighted and weighted networks and apply it to a faculty hiring dataset, discovering a meaningful hierarchy of communities among US business schools.

研究动机与目标

  • 解决通过调查收集数据时引入采样偏差和噪声的网络中社区检测的挑战。
  • 构建一个通用的统计框架,将边采样概率建模为个体偏好和社区归属的函数。
  • 在调查采样机制下实现一致的社区检测,克服标准社区检测方法的局限性。
  • 提出一种参数化模型——提名随机块模型,用于有向网络,具备可解释的参数和高效估计。
  • 为所提出的框架下谱聚类和矩方法估计提供理论一致性保证。

提出的方法

  • 将边采样概率建模为个体层面偏好和社区分配的函数,捕捉调查响应行为中的异质性。
  • 制定一个通用的网络采样模型,将随机块模型扩展至包含调查设计中的采样机制。
  • 在所提出的模型下对观测到的采样网络应用谱聚类,并在温和正则性条件下证明其一致性。
  • 将提名随机块模型作为有向网络的特例提出,实现可解释的参数估计。
  • 使用矩方法高效估计模型参数,并提供理论一致性保证。
  • 将谱聚类与矩方法整合到统一框架中,用于调查采样网络中的社区检测。

实验结果

研究问题

  • RQ1当网络数据通过调查采样收集时,如何可靠地执行社区检测,该过程会引入偏差和缺失边?
  • RQ2是否存在一种统计上合理的方法,用于建模依赖于个体偏好和社区结构的边采样概率?
  • RQ3谱聚类在这一类通用网络采样模型下是否仍能保持一致性?
  • RQ4如何构建有向网络的参数化模型,以实现参数的有意义解释和高效估计?
  • RQ5在调查采样机制下,所提出的估计方法具有怎样的理论一致性性质?

主要发现

  • 所提出的模型在一般调查采样机制下,可通过谱聚类实现一致的社区检测。
  • 提名随机块模型为有向网络提供了一个参数化、可解释的框架,支持高效的矩方法估计。
  • 在所提出的采样模型下,提名随机块模型的矩方法估计器是一致的。
  • 模拟研究结果表明,在各种采样方案下,该方法在无权和加权网络中均能准确恢复社区结构。
  • 在一所高校招聘数据集上的应用揭示了美国商学院之间有意义的分层结构,验证了该模型的实证有效性。
  • 理论分析证实,在所提出的模型假设下,谱聚类和矩方法估计均具有一致性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。