Skip to main content
QUICK REVIEW

[论文解读] Estimation and clustering in a semiparametric Poisson process stochastic block model for longitudinal networks

Catherine Matias, Tabea Rebafka|arXiv (Cornell University)|Dec 21, 2015
Bayesian Methods and Mixture Models参考文献 36被引用 7
一句话总结

该论文提出了一种纵向交互网络的半参数泊松过程随机块模型,其中使用核密度或直方图估计器在变分EM框架内对特定群体的交互强度进行非参数建模。主要贡献是一种一致估计程序,结合集成分类似然准则(ICL)以选择群体数量,其理论一致性结果在强度和群体估计方面均得到支持。

ABSTRACT

In this work, we introduce a Poisson process stochastic block model for recurrent interaction events, where each individual belongs to a latent group and interactions between two individuals follow a conditional inhomogeneous Poisson process whose intensity is driven by the individuals’ latent groups. The model is semiparametric as the intensities per group pair are modeled in a nonparametric way. First an identifiability result on the weights of the latent groups and the nonparametric intensities is established. Then we propose an estimation procedure, relying on a semiparametric version of a variational expectation-maximization algorithm. Two different versions of the method are proposed, using either histogram-type (with an adaptive choice of the partition size) or kernel intensity estimators. We also propose an integrated classification likelihood criterion to select the number of latent groups. Asymptotic consistency results are then explored, both for the estimators of the cumulative intensities per group pair and for the kernel procedures that estimate the intensities per group pair. Finally, we carry out synthetic experiments and analyse several real datasets to illustrate the strengths and weaknesses of our approach.

研究动机与目标

  • 使用具有群体特异性条件强度的随机块模型对纵向网络中的重复交互事件进行建模。
  • 解决动态交互数据中群体结构未知且强度函数非参数化的问题。
  • 开发一种一致估计程序,联合推断潜在群体与交互强度。
  • 提供一种模型选择准则,以确定网络中潜在群体的数量。

提出的方法

  • 将交互建模为条件非齐次泊松过程,其强度依赖于交互个体的潜在群体。
  • 采用半参数方法,通过核密度或直方图估计器对群体对之间的强度进行非参数建模。
  • 使用变分期望-最大化算法,联合估计群体归属与强度函数。
  • 在非参数强度估计中引入自适应带宽或划分大小选择。
  • 应用集成分类似然(ICL)准则以选择最优潜在群体数量。
  • 证明了累积强度估计器与基于核的强度程序在理论上的的一致性。

实验结果

研究问题

  • RQ1如何在纵向事件数据网络中,对潜在群体之间的非参数交互强度实现一致估计?
  • RQ2该模型中群体权重与非参数强度函数的可辨识性条件是什么?
  • RQ3在具有复杂强度结构的半参数设定下,如何有效选择潜在群体的数量?
  • RQ4所提出的累积强度与个体强度函数估计量的渐近行为如何?
  • RQ5在有限样本下,核方法与直方图估计方法的性能表现有何差异?

主要发现

  • 在较弱的正则性条件下,模型可实现群体权重与非参数强度函数的可辨识性。
  • 所提出的结合非参数强度估计的变分EM算法,对累积强度估计器与基于核的强度估计均表现出一致性。
  • 集成分类类似然(ICL)准则在模拟研究中成功识别出真实的潜在群体数量。
  • 自适应直方图与核估计方法在有限样本下表现优异,其中核方法提供更平滑的强度估计。
  • 在观测窗口与群体规模逐渐增大的条件下,理论一致性已证明适用于累积强度估计器与基于核的强度程序。
  • 在合成数据与真实数据集上的实证分析证实,该模型能够有效恢复有意义的群体结构与动态交互模式。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。