Skip to main content
QUICK REVIEW

[논문 리뷰] Refining Approximating Betweenness Centrality Based on Samplings

Shiyu Ji, Zenghui Yan|arXiv (Cornell University)|2016. 08. 16.
Complex Network Analysis Techniques참고 문헌 17인용 수 6
한 줄 요약

이 논문은 Bader의 샘플링 기반 중심성 중심성(BC) 근사 알고리즘 재검토하여 그 경계에 이론적 결함을 밝혀내고, 더 날카운 경량 확률 보장을 갖는 개선된 샘플링 방법을 제안한다. 정점 쌍에 기반한 적응형 알고리즘을 도입하여 이전 작업보다 낮은 근사 요소, 높은 성공 확률, 더 적은 샘플 수를 달성하였으며, 실제 네트워크에서 검증된 결과 최대 7%의 요소 차이 향상이 이루어졌다.

ABSTRACT

Betweenness Centrality (BC) is an important measure used widely in complex network analysis, such as social network, web page search, etc. Computing the exact BC values is highly time consuming. Currently the fastest exact BC determining algorithm is given by Brandes, taking $O(nm)$ time for unweighted graphs and $O(nm+n^2\log n)$ time for weighted graphs, where $n$ is the number of vertices and $m$ is the number of edges in the graph. Due to the extreme difficulty of reducing the time complexity of exact BC determining problem, many researchers have considered the possibility of any satisfactory BC approximation algorithms, especially those based on samplings. Bader et al. give the currently best BC approximation algorithm, with a high probability to successfully estimate the BC of one vertex within a factor of $1/\varepsilon$ using $\varepsilon t$ samples, where $t$ is the ratio between $n^2$ and the BC value of the vertex. However, some of the algorithmic parameters in Bader's work are not yet tightly bounded, leaving some space to improve their algorithm. In this project, we revisit Bader's algorithm, give more tight bounds on the estimations, and improve the performance of the approximation algorithm, i.e., fewer samples and lower factors. On the other hand, Riondato proposes a non-adaptive BC approximation method with samplings on shortest paths. We investigate the possibility to give an adaptive algorithm with samplings on vertices. With rigorous reasoning, we show that this algorithm can also give bounded approximation factors and number of samplings. To evaluate our results, we also conduct extensive experiments using real-world datasets, and find that both our algorithms can achieve acceptable performance. We verify that our work can improve the current BC approximation algorithm with better parameters, i.e., higher successful probabilities, lower factors and fewer samples.

연구 동기 및 목표

  • Bader의 BC 근사 알고리즘의 이론적 간극, 특히 잘못된 확률 경계와 증명되지 않은 보조정리들을 해결하기 위해.
  • 샘플 수와 근사 요소를 줄임으로써 BC 근사의 성능을 향상시키기 위해.
  • 근사 오차가 제한된 새로운 적응형 샘플링 전략을 개발하기 위해.
  • 실제 그래프 데이터셋을 사용하여 새로운 알고리즘을 철저히 분석하고 검증하기 위해.
  • 기존 방법보다 더 날카운 이론적 경계를 확립하기 위해 근사 요소와 샘플링 복잡도에 대해.

제안 방법

  • Bader의 알고리즘을 재검토하고, 특히 보조정리 3에서의 잘못된 확률 추론을 수정하기 위해 새로운 확률 모델을 개발한다.
  • Bader의 전략과 Riondato의 샘플링 전략을 융합한 방식을 사용한 개선된 Bader 스타일 알고리즘을 제안한다.
  • 동적 임계값에 기반해 정점 쌍을 선택함으로써 추정 정확도를 향상시키는 적응형 샘플링 방법을 도입한다.
  • 근사 요소가 $ b = t^{1/3}/S $로 제한됨을 보여주는 이론적 경계를 유도한다. 여기서 $ S $는 샘플 수이고 $ t $는 $ n^2 $과 BC 값의 비율이다.
  • 샘플링 루프에서 임계값 기반 종료 조건을 적용하며, $ c $ 값이 클수록 더 많은 샘플과 더 날카운 경계를 허용한다.
  • 실제 그래프에서 광범위한 실험을 통해 요소 차이와 샘플 수를 비교하여 성능을 검증한다.

실험 결과

연구 질문

  • RQ1Bader의 BC 근사 알고리즘에서 이론적 경계를 더 날카우게 조정하여 정확도를 높이고 샘플링 요구량을 줄일 수 있는가?
  • RQ2Bader의 원본 증명에서 잘못된 확률 추론이 미치는 영향은 무엇이며, 이를 어떻게 수정할 수 있는가?
  • RQ3정점 쌍에 기반한 적응형 샘플링 전략이 균일 샘플링보다 더 낮은 근사 요소와 더 적은 샘플 수로 더 나은 성능을 낼 수 있는가?
  • RQ4샘플 수와 BC 근사에서 요소의 역수 사이에 이론적 관계가 존재하는가?
  • RQ5제안된 알고리즘이 기존 방법에 비해 실제 희박한 네트워크에서 어떻게 성능을 발휘하는가?

주요 결과

  • 검토된 그래프 전반에서 개선된 Bader 알고리즘은 800개 미만의 샘플로도 요소 차이를 7% 이내로 줄였다.
  • 요소 차이의 역수는 샘플 수와 약선형 비례함을 확인하여 이론적 예측을 뒷받침한다.
  • 적응형 쌍 기반 샘플링 알고리즘이 요소 차이를 5% 미만으로 낮추어 개선된 Bader 방법보다 정확도에서 뛰어난 성능을 보였다.
  • 이론적 분석을 통해 근사 요소가 $ b = t^{1/3}/S $로 제한됨을 확인하였으며, Bader의 원본 작업보다 더 날카운 경계를 확보하였다.
  • 샘플 수는 임계값 매개변수 $ c $와 거의 선형적으로 증가함을 확인하여 예측 가능하고 확장 가능한 행동을 보였다.
  • 실험 결과 두 알고리즘이 이전의 샘플링 기반 BC 근사 방법보다 더 높은 성공 확률과 더 낮은 요소를 더 적은 샘플 수로 달성함을 확인하였다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.