Skip to main content
QUICK REVIEW

[論文レビュー] Predicting Graph Categories from Structural Properties

James P. Canning, Emma E. Ingram|arXiv (Cornell University)|May 7, 2018
Complex Network Analysis Techniques参考文献 47被引用数 6
ひとこと要約

本稿では、構造的性質のみを用いて複雑ネットワークのドメインカテゴリを予測する手法を提案し、実際のネットワークと合成ネットワークを組み合わせたデータセットでランダムフォレスト分類器を用いて96.6%の精度を達成した。異なる構造的指紋が高精度分類を可能にすることを示し、合成グラフはその生成モデルに基づいてほぼ完璧に分類可能であることがわかった。

ABSTRACT

This paper has been withdrawn from arXiv.org due to a disagreement among the authors related to several peer-review comments received prior to submission on arXiv.org. Even though the current version of this paper is withdrawn, there was no disagreement between authors on the novel work in this paper. One specific issue was the discussion of related work by Ikehara \& Clauset (found on page 8 of the previously posted version). Peer-review comments on a similar version made ALL authors aware that the discussion misrepresented their work prior to submission to arXiv.org. However, some authors choose to post to arXiv a minimally updated version without the consent of all authors or properly addressing this attribution issue. ================ Original Paper Abstract: Complex networks are often categorized according to the underlying phenomena that they represent such as molecular interactions, re-tweets, and brain activity. In this work, we investigate the problem of predicting the category (domain) of arbitrary networks. This includes complex networks from different domains as well as synthetically generated graphs from five different network models. A classification accuracy of $96.6\%$ is achieved using a random forest classifier with both real and synthetic networks. This work makes two important findings. First, our results indicate that complex networks from various domains have distinct structural properties that allow us to predict with high accuracy the category of a new previously unseen network. Second, synthetic graphs are trivial to classify as the classification model can predict with near-certainty the network model used to generate it. Overall, the results demonstrate that networks drawn from different domains (and network models) are trivial to distinguish using only a handful of simple structural properties.

研究の動機と目的

  • 異なるドメインからの複雑ネットワークが、構造的性質のみに基づいて分類可能かどうかを調査すること。
  • 機械学習モデルの性能を、構造的性質のみを用いて実世界ネットワークと合成ネットワークを区別できるかを評価すること。
  • 合成ネットワークモデルが、正確な分類を可能にする構造的インプリントを残すかどうかを特定すること。
  • 多様なネットワークドメインおよびモデルにわたる構造的特徴の一般化可能性を評価すること。

提案手法

  • 各ネットワークから抽出された14個の単純な構造的性質に基づいてランダムフォレスト分類器を訓練する。
  • データセットには、5つのドメインからの実ネットワークと、5つの異なるネットワークモデルから生成された合成グラフが含まれる。
  • 特徴量は正規化され、実ネットワークと合成ネットワークの組み合わせデータセット上で分類器を訓練する。
  • モデルの性能は10分割交差検証を用いて評価され、妥当性と一般化能力を確保する。
  • 分類器は、以前に見られなかったネットワークを用いて予測精度を評価する。
  • 研究は構造的特徴に限定され、ドメイン固有の情報や意味的情報は一切使用しない。

実験結果

リサーチクエスチョン

  • RQ1異なるドメインからの複雑ネットワークは、構造的性質のみに基づいて正確に分類可能か?
  • RQ2機械学習モデルは、構造的性質のみを用いて実世界ネットワークと合成ネットワークをどの程度正確に区別できるか?
  • RQ3合成ネットワークモデルは、同定を可能にする独自の構造的サインをどの程度残すか?
  • RQ4多様なドメインにわたるネットワークカテゴリの予測において、特に予測に寄与する構造的特徴は何か?

主な発見

  • ランダムフォレスト分類器は、構造的性質のみを用いてネットワークカテゴリを区別する際、96.6%の分類精度を達成した。
  • 合成グラフはほぼ完全に分類可能であり、各ネットワークモデルが一意で識別可能な構造的パターンを生成していることが示された。
  • 異なるドメインからのネットワークが固有の構造的指紋を示すことが確認され、高精度な分類に十分であることがわかった。
  • 本研究は、単純な構造的特徴が、実ネットワークおよび合成グラフを含む多様なネットワークタイプにおいて、顕著な識別能力を示すことを示した。
  • 高い精度は、構造的性質のみでネットワークカテゴリを区別可能であることを示唆しており、ドメイン固有の知識がなくても十分である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。