Skip to main content
QUICK REVIEW

[论文解读] Latent Network Models to Account for Noisy, Multiply-Reported Social Network Data

Caterina De Bacco, Martina Contisciani|arXiv (Cornell University)|Dec 21, 2021
Complex Network Analysis Techniques被引用 5
一句话总结

本文提出了一种概率潜在网络模型,通过从多个个体报告中估计未观测到的网络结构,来处理多重报告社会网络数据中的噪声和互惠性。利用变分推断,该模型捕捉了报告者特定的报告偏差和相互报告倾向(互惠性),揭示了印度和尼加拉瓜村庄网络中显著的互惠效应,而这些效应在标准聚合方法中被忽略。

ABSTRACT

Social network data are often constructed by incorporating reports from multiple individuals. However, it is not obvious how to reconcile discordant responses from individuals. There may be particular risks with multiply-reported data if people's responses reflect normative expectations -- such as an expectation of balanced, reciprocal relationships. Here, we propose a probabilistic model that incorporates ties reported by multiple individuals to estimate the unobserved network structure. In addition to estimating a parameter for each reporter that is related to their tendency of over- or under-reporting relationships, the model explicitly incorporates a term for ``mutuality,'' the tendency to report ties in both directions involving the same alter. Our model's algorithmic implementation is based on variational inference, which makes it efficient and scalable to large systems. We apply our model to data from 75 Indian villages collected with a name-generator design, and a Nicaraguan community collected with a roster-based design. We observe strong evidence of ``mutuality'' in both datasets, and find that this value varies by relationship type. Consequently, our model estimates networks with reciprocity values that are substantially different than those resulting from standard deterministic aggregation approaches, demonstrating the need to consider such issues when gathering, constructing, and analysing survey-based network data.

研究动机与目标

  • 解决在调查数据中调和多个个体报告的不一致社会网络报告的挑战。
  • 对报告者特定的过度或不足报告关系行为进行建模。
  • 明确考虑多个报告者之间的互惠性——即报告互惠关系的倾向。
  • 开发一种可扩展的、高效的推断方法,用于从噪声大、多重报告的数据中重建大规模网络。
  • 使用真实世界村庄数据集,展示互惠性和报告偏差对网络结构的影响。

提出的方法

  • 该模型采用概率生成框架,从多个噪声报告中推断潜在的真实网络。
  • 引入报告者特定的参数,以建模个体在过度或不足报告关系方面的倾向。
  • 设置独立的互惠性参数,以捕捉报告者报告互惠关系的倾向,即使仅有一方被报告。
  • 该模型采用变分推断,实现高效且可扩展的后验近似,支持在大规模网络上的应用。
  • 似然函数结合了报告者特定的报告概率和互惠报告项,联合估计网络结构和报告偏差。
  • 算法通过均值场变分推断迭代更新参数,以最大化边缘似然的下界。

实验结果

研究问题

  • RQ1个体报告偏差在多大程度上影响聚合社会网络数据的准确性?
  • RQ2在多重报告网络数据中,互惠性——即互惠报告关系——的程度如何?
  • RQ3不同关系类型在自我报告网络中对互惠性强弱有何影响?
  • RQ4推断出的潜在网络结构与通过标准确定性聚合方法获得的结构有何不同?
  • RQ5一种可扩展的概率模型能否有效校正基于调查的数据中的噪声和互惠性?

主要发现

  • 在印度村庄和尼加拉瓜社区数据集中均发现强烈的互惠性证据,表明存在系统性的互惠报告行为。
  • 互惠性在不同关系类型间存在显著差异,表明报告行为在不同社会角色中并非一致。
  • 该模型估计的互惠性值与通过标准确定性聚合方法获得的结果存在显著差异,凸显了传统方法中的偏差风险。
  • 报告者特定的偏差参数揭示了系统性的过度或不足报告倾向,且在个体间存在差异。
  • 变分推断的实现支持在大规模网络数据集上实现高效且可扩展的推断,具备实际应用潜力。
  • 结果表明,忽略互惠性和报告偏差会导致网络结构失真,尤其在主观或基于感知的网络数据中。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。