Skip to main content
QUICK REVIEW

[论文解读] A maximum entropy approach to separating noise from signal in bimodal affiliation networks

Navid Dianati|arXiv (Cornell University)|Jul 6, 2016
Complex Network Analysis Techniques参考文献 12被引用 4
一句话总结

本文提出了一种基于最大熵的零模型,通过保留双模隶属网络中两层节点的度序列(实体频率和集合大小),实现从噪声中分离信号。该方法使用对泊松二项分布的改进正态近似,计算共现权重的p值,从而在无需蒙特卡洛采样的情况下实现高效的显著性检验。与权重阈值法相比,该方法在揭示模块化结构方面表现更优,例如在第110届美国参议院共赞助网络中清晰呈现了两党制分歧。

ABSTRACT

In practice, many empirical networks, including co-authorship and collocation networks are unimodal projections of a bipartite data structure where one layer represents entities, the second layer consists of a number of sets representing affiliations, attributes, groups, etc., and an inter-layer link indicates membership of an entity in a set. The edge weight in the unimodal projection, which we refer to as a co-occurrence network, counts the number of sets to which both end-nodes are linked. Interpreting such dense networks requires statistical analysis that takes into account the bipartite structure of the underlying data. Here we develop a statistical significance metric for such networks based on a maximum entropy null model which preserves both the frequency sequence of the individuals/entities and the size sequence of the sets. Solving the maximum entropy problem is reduced to solving a system of nonlinear equations for which fast algorithms exist, thus eliminating the need for expensive Monte-Carlo sampling techniques. We use this metric to prune and visualize a number of empirical networks.

研究动机与目标

  • 为解决从双模数据导出的密集共现网络中识别统计显著边的挑战。
  • 开发一种计算高效的蒙特卡洛方法替代方案,用于双模网络的零模型。
  • 以统计上合理的方式同时保留实体的频率序列和集合的大小序列。
  • 实现对噪声边的稳健修剪,同时保留多尺度网络结构。
  • 展示在检测社区结构(如立法网络中的两党制)方面的改进效果。

提出的方法

  • 构建一个最大熵零模型,以保留双模网络中两层节点的期望度序列。
  • 将最大熵问题的求解转化为求解非线性方程组,避免昂贵的蒙特卡洛采样。
  • 使用泊松二项分布来建模在零模型下任意两个实体之间的期望共现频率。
  • 应用泊松二项分布累积分布函数的改进正态近似,以实现p值的快速计算。
  • 通过检验统计量 $-\log(\pi_{ij})$ 计算显著性,其中 $\pi_{ij}$ 是观察到的共现权重 $w_{ij}$ 的右尾p值。
  • 通过仅保留显著性值高于阈值的边来修剪网络,从而揭示网络的骨干结构。

实验结果

研究问题

  • RQ1如何在保留底层结构约束的前提下,识别双模隶属网络中的统计显著共现边?
  • RQ2最大熵零模型是否能在检测有意义的网络结构方面优于现有的基于随机化的方法?
  • RQ3所提出的方法在多大程度上改善了对立法网络中模块化结构(如两党制)的检测?
  • RQ4与朴素的权重阈值法相比,显著性过滤器在巨型连通分量大小和模块化程度方面的表现如何?
  • RQ5该方法是否能无需依赖计算成本高昂的蒙特卡洛采样而实现高效计算?

主要发现

  • 所提出的显著性过滤器在第110届美国参议院共赞助网络中揭示了高度模块化的结构,清晰反映了民主党和共和党的党派分歧。
  • 在网络密度为2时,使用显著性过滤器修剪后,巨型连通分量包含约80%的参议员,而权重阈值法下的连通分量则小得多。
  • 在所有截断水平下,使用显著性过滤器修剪后的网络模块化程度始终显著高于权重阈值法。
  • 该方法通过求解非线性方程组而非依赖蒙特卡洛采样来估计零分布,实现了高计算效率。
  • 对泊松二项分布CDF的改进正态近似使得p值计算快速且准确,使该方法可扩展至大规模网络。
  • 通过过滤器识别出的共现网络骨干结构呈现出清晰且可解释的社区结构,而在原始密集网络中这一结构则被掩盖。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。