Skip to main content
QUICK REVIEW

[論文レビュー] Counting Motifs with Graph Sampling

Jason M. Klusowski, Yihong Wu|arXiv (Cornell University)|Feb 21, 2018
Complex Network Analysis Techniques参考文献 29被引用数 3
ひとこと要約

本稿は、大規模ネットワークにおけるモチーフ(誘導部分グラフ)の統計的推定を、部分グラフサンプリングおよび近隣領域サンプリングを用いて行う。乗法的誤差でモチーフカウントを推定するための最適なサンプリング比を確立し、近隣領域サンプリングが部分グラフサンプリングを著しく上回ることを示しており、その結果はグラフのトポロジーに依存せず、モチーフのサイズとグラフの次数にのみ依存する。

ABSTRACT

Applied researchers often construct a network from a random sample of nodes in order to infer properties of the parent network. Two of the most widely used sampling schemes are subgraph sampling, where we sample each vertex independently with probability $p$ and observe the subgraph induced by the sampled vertices, and neighborhood sampling, where we additionally observe the edges between the sampled vertices and their neighbors. In this paper, we study the problem of estimating the number of motifs as induced subgraphs under both models from a statistical perspective. We show that: for any connected $h$ on $k$ vertices, to estimate $s=\mathsf{s}(h,G)$, the number of copies of $h$ in the parent graph $G$ of maximum degree $d$, with a multiplicative error of $ε$, (a) For subgraph sampling, the optimal sampling ratio $p$ is $Θ_{k}(\max\{ (sε^2)^{-\frac{1}{k}}, \; \frac{d^{k-1}}{sε^{2}} \})$, achieved by Horvitz-Thompson type of estimators. (b) For neighborhood sampling, we propose a family of estimators, encompassing and outperforming the Horvitz-Thompson estimator and achieving the sampling ratio $O_{k}(\min\{ (\frac{d}{sε^2})^{\frac{1}{k-1}}, \; \sqrt{\frac{d^{k-2}}{sε^2}}\})$. This is shown to be optimal for all motifs with at most $4$ vertices and cliques of all sizes. The matching minimax lower bounds are established using certain algebraic properties of subgraph counts. These results quantify how much more informative neighborhood sampling is than subgraph sampling, as empirically verified by experiments on both synthetic and real-world data. We also address the issue of adaptation to the unknown maximum degree, and study specific problems for parent graphs with additional structures, e.g., trees or planar graphs.

研究の動機と目的

  • サンプリングされた部分グラフしか入手できない状況下で、大規模ネットワークにおけるモチーフカウントを推定する統計的課題に取り組む。
  • 現実的なサンプリング制約下で、部分グラフサンプリングと近隣領域サンプリングの間でモチーフカウントの効率性を比較する。
  • 任意の連結モチーフ(サイズk)に対して推定誤差を最小化する最適なサンプリング比を導出する。
  • 両方のサンプリングモデル下で、普遍的に最適な推定器(特にHorvitz-Thompson型推定器)を構築する。
  • 最小最大下界を確立し、小規模なモチーフおよびクリークに対して提案された推定器の最適性を示す。

提案手法

  • 部分グラフサンプリングに対して、すべての連結モチーフに対して普遍的に最適であるHorvitz-Thompson型推定器を提案する。
  • Horvitz-Thompson推定器を一般化し、性能を上回る推定器の族を近隣領域サンプリングに対して導入する。
  • 部分グラフカウントの代数的性質と最小最大リスク解析を用いて、最適なサンプリング比を導出する。
  • 組合せ的および代数的技法を用いて一致する最小最大下界を確立し、提案されたサンプリング戦略の最適性を証明する。
  • データ駆動型のサンプリング比選択を用いて、親グラフの最大次数dが未知である場合の適応性を分析する。
  • 木構造や平面グラフなどの特別なグラフ構造を検討し、構造的制約下でのサンプリング効率を高める。

実験結果

リサーチクエスチョン

  • RQ1部分グラフサンプリングにおいて、乗法的誤差εでモチーフカウントを推定するための最適なサンプリング比pは何か?
  • RQ2モチーフカウントにおいて、近隣領域サンプリングは部分グラフサンプリングに比べてどのように推定効率を向上させるか?
  • RQ3両方のサンプリングモデル下で、すべての連結モチーフに対して普遍的に最適な推定器を構築できるか?
  • RQ4部分グラフサンプリングおよび近隣領域サンプリング下でのモチーフカウントに対する最小最大下界は何か?
  • RQ5親グラフの最大次数dが未知である場合、どのようにしてサンプリング比を適応的に選択できるか?

主な発見

  • 部分グラフサンプリングでは、最適なサンプリング比はΘk(max{(sε²)^(-1/k), d^{k-1}/(sε²)})であり、モチーフのトポロジーに依存しない。
  • Horvitz-Thompson推定器は、部分グラフサンプリング下で、すべての連結モチーフに対して普遍的に最適である。
  • 近隣領域サンプリングでは、提案された推定器がOk(min{(d/(sε²))^{1/(k-1)}, √(d^{k-2}/(sε²))})のサンプリング比を達成し、再びモチーフ構造に依存しない。
  • 近隣領域サンプリング推定器は、4頂点以下のすべてのモチーフおよび任意サイズのクリークに対して最適である。
  • 最小最大下界が上界と一致しており、提案されたサンプリング戦略の最適性が証明されている。
  • 合成データおよび実世界データにおける実験により、近隣領域サンプリングが部分グラフサンプリングよりも著しく高い推定精度を達成することが確認された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。