Skip to main content
QUICK REVIEW

[論文レビュー] Characterizing Directed and Undirected Networks via Multidimensional Walks with Jumps

Fabrício Murai, Bruno Ribeiro|arXiv (Cornell University)|Mar 23, 2017
Complex Network Analysis Techniques参考文献 21被引用数 4
ひとこと要約

本稿では、強連結成分に閉じ込められるのを防ぐためにランダムジャンプを併用した複数の連携するウォークャーを用いる、新しいランダムウォークベースの手法であるDirected Unbiased Frontier Sampling(DUFS)を提案する。DUFSはウォークャーの軌跡からリアルタイムで無向グラフを構築し、ノードラベル分布の改善された最小分散不偏推定量を提供する。これは、出次数分布の先頭部の推定において既存手法を著しく上回り、テール部の推定精度においても同等またはそれを上回る。

ABSTRACT

Estimating distributions of node characteristics (labels) such as number of connections or citizenship of users in a social network via edge and node sampling is a vital part of the study of complex networks. Due to its low cost, sampling via a random walk (RW) has been proposed as an attractive solution to this task. Most RW methods assume either that the network is undirected or that walkers can traverse edges regardless of their direction. Some RW methods have been designed for directed networks where edges coming into a node are not directly observable. In this work, we propose Directed Unbiased Frontier Sampling (DUFS), a sampling method based on a large number of coordinated walkers, each starting from a node chosen uniformly at random. It is applicable to directed networks with invisible incoming edges because it constructs, in real-time, an undirected graph consistent with the walkers trajectories, and due to the use of random jumps which prevent walkers from being trapped. DUFS generalizes previous RW methods and is suited for undirected networks and to directed networks regardless of in-edges visibility. We also propose an improved estimator of node label distributions that combines information from the initial walker locations with subsequent RW observations. We evaluate DUFS, compare it to other RW methods, investigate the impact of its parameters on estimation accuracy and provide practical guidelines for choosing them. In estimating out-degree distributions, DUFS yields significantly better estimates of the head of the distribution than other methods, while matching or exceeding estimation accuracy of the tail. Last, we show that DUFS outperforms uniform node sampling when estimating distributions of node labels of the top 10% largest degree nodes, even when sampling a node uniformly has the same cost as RW steps.

研究の動機と目的

  • 入力エッジが見えない有向ネットワークにおいて、標準的なランダムウォークがグラフ全体を探索できないという課題に対処すること。
  • 既存のランダムウォークサンプリング手法を、隠れたインバウンドエッジを含む有向および無向ネットワークの両方で動作可能に一般化すること。
  • 初期ウォークャーの位置とその後のウォーク観測値を統合する、ノードラベル分布の改善された推定量を開発すること。
  • 推定精度を最適化するためのDUFSパラメータの選定に関する実用的ガイドラインを提供すること。
  • 単一の均一ノードクエリのコストがランダムウォークステップのコストと同等である場合でも、DUFSが均一ノードサンプリングを上回ることを示すこと。

提案手法

  • DUFSは、均一にランダムに選ばれたノードから出発する複数の連携するランダムウォークャーを採用する。
  • 各ウォークャーは、強連結成分に閉じ込められるのを防ぐためにランダムジャンプを伴うランダムウォークを実行する。
  • 本手法は、すべてのウォークャーの軌跡から動的に無向グラフを構築し、入力エッジが非表示の有向ネットワークでも不偏なサンプリングを可能にする。
  • 初期ウォークャー位置とその後のウォーク観測値の情報を統合する新しい推定量を導入し、十分統計量 $ n_i + m_i $ を用いる。
  • Fisher-Neymanの因数分解定理と $ n_i + m_i $ の完全性を用いて、この推定量が最小分散不偏推定量(MVUE)であることを理論的に証明した。
  • 本手法は、フロンティアサンプリング(FS)と有向不偏ランダムウォーク(DURW)を一般化し、両者の長所を統合して、無向および有向ネットワークの両方で機能する。

実験結果

リサーチクエスチョン

  • RQ1入力エッジが観測不能な有向ネットワークにおいても、効果的に動作するランダムウォークベースのサンプリング手法を設計できるか?
  • RQ2強連結成分に閉じ込められるのを防ぎつつ、ノードラベル分布の不偏推定を維持するようにランダムウォークをどのように変更できるか?
  • RQ3このようなサンプリングフレームワークにおいて、ノードラベル分布の最小分散不偏推定を達成する推定量の構造は何か?
  • RQ4サンプリングプロセスのパラメータ(例:ウォークャー数、ジャンプ確率)が推定精度に与える影響は何か?
  • RQ5サンプリングコストが同等である場合でも、DUFSは高次数ノードのラベル分布推定において、均一ノードサンプリングを上回る精度を示すか?

主な発見

  • DUFSは、他のランダムウォーク手法と比較して、出次数分布の先頭部の推定において著しく優れた結果を得た。
  • 既存手法と比較して、出次数分布のテール部の推定精度は同等またはそれを上回った。
  • 提案された推定量は、十分統計量と完全統計量に基づく理論的保証を有し、最小分散不偏推定量(MVUE)であることが証明された。
  • 単一の均一ノードクエリのコストがランダムウォークステップのコストと同等である場合でも、DUFSは上位10%の最大次数ノードのラベル分布推定において、均一ノードサンプリングを上回った。
  • 本手法は有向および無向ネットワークの両方で頑健であり、フロンティアサンプリングやDURWといった先行手法を一般化した。
  • パラメータ感度分析により、推定精度を最適化するためのウォークャー数とジャンプ確率のチューニングに関する実用的ガイドラインが得られた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。