Skip to main content
QUICK REVIEW

[논문 리뷰] Effective Deterministic Initialization for $k$-Means-Like Methods via Local Density Peaks Searching

Fengfu Li, Hong Qiao|arXiv (Cornell University)|2016. 11. 21.
Statistical Methods and Inference참고 문헌 28인용 수 4
한 줄 요약

이 논문은 $k$-means 유사 클러스터링을 위한 결정론적 초기화 프레임워크인 국소 밀도 피크 탐색(LDPS)을 제안한다. 이는 클러스터 수 $k$를 자동으로 추정하고, 높은 품질의 초기 중심을 선택하며, 이국적 데이터를 탐지하고, 비유클리드 거리도 지원한다. 국소 밀도와 새로운 국소 독특성 지수(LDI)를 활용함으로써 LDPS-means와 LDPS-medoids는 특히 큰 $k$와 비구형 데이터에서 표준 $k$-means와 $k$-medoids보다 훨씬 적은 반복 횟수로 더 높은 정확도를 달성하는 뛰어난 클러스터링 성능을 보인다.

ABSTRACT

The $k$-means clustering algorithm is popular but has the following main drawbacks: 1) the number of clusters, $k$, needs to be provided by the user in advance, 2) it can easily reach local minima with randomly selected initial centers, 3) it is sensitive to outliers, and 4) it can only deal with well separated hyperspherical clusters. In this paper, we propose a Local Density Peaks Searching (LDPS) initialization framework to address these issues. The LDPS framework includes two basic components: one of them is the local density that characterizes the density distribution of a data set, and the other is the local distinctiveness index (LDI) which we introduce to characterize how distinctive a data point is compared with its neighbors. Based on these two components, we search for the local density peaks which are characterized with high local densities and high LDIs to deal with 1) and 2). Moreover, we detect outliers characterized with low local densities but high LDIs, and exclude them out before clustering begins. Finally, we apply the LDPS initialization framework to $k$-medoids, which is a variant of $k$-means and chooses data samples as centers, with diverse similarity measures other than the Euclidean distance to fix the last drawback of $k$-means. Combining the LDPS initialization framework with $k$-means and $k$-medoids, we obtain two novel clustering methods called LDPS-means and LDPS-medoids, respectively. Experiments on synthetic data sets verify the effectiveness of the proposed methods, especially when the ground truth of the cluster number $k$ is large. Further, experiments on several real world data sets, Handwritten Pendigits, Coil-20, Coil-100 and Olivetti Face Database, illustrate that our methods give a superior performance than the analogous approaches on both estimating $k$ and unsupervised object categorization.

연구 동기 및 목표

  • 표준 $k$-means의 한계, 즉 초기값에 대한 민감성, 비구형 클러스터에서의 열악한 성능, 그리고 $k$를 사전에 지정해야 하는 요구사항을 해결하기 위해.
  • 기하학적으로 진정한 클러스터 중심에 가까운 높은 품질의 클러스터 중심을 선택하는 결정론적 초기화 방법을 개발하기 위해.
  • 국소 밀도와 새로운 국소 독특성 지수(LDI)를 사용하여 클러스터 수 $k$를 사전 지식 없이 자동으로 추정하기 위해.
  • 이국적 데이터를 클러스터링 이전에 탐지하고 제거하여 강건성을 향상시키기 위해.
  • 복잡한 데이터 분포에서 더 나은 성능을 내기 위해 다양체 기반 유사도 측정법을 사용하는 $k$-medoids로 프레임워크를 확장하기 위해.

제안 방법

  • 이 방법은 데이터 포인트의 이웃 기반으로 밀도 분포를 특성화하는 국소 밀도 측정법을 도입한다.
  • 이웃에 비해 얼마나 독특한지 수량화하는 국소 독특성 지수(LDI)를 정의한다. 이는 높은 밀도이면서 동시에 고립된 포인트를 선호한다.
  • 국소 밀도가 높고 LDI도 높은 포인트, 즉 국소 밀도 피크로 간주되며, 이들이 초기 클러스터 중심으로 선택된다.
  • 국소 밀도는 낮지만 LDI는 높은 포인트로 이국적 데이터를 탐지하고, 클러스터링 이전에 제거한다.
  • LDPS 프레임워크는 $k$-means와 $k$-medoids에 통합되어 각각 LDPS-means와 LDPS-medoids를 도출한다.
  • 다양체 구조를 가진 데이터의 경우, 그래프 기반 거리와 CW-SSIM 지수를 비유사도 측정법으로 사용하여 $k$-medoids에서 활용한다.

실험 결과

연구 질문

  • RQ1결정론적 초기화 방법이 랜덤 시드 선택에 민감한 $k$-means의 성능을 향상시킬 수 있는가?
  • RQ2국소 밀도와 LDI를 사용하여 사전 지식 없이 클러스터 수 $k$를 자동으로 추정할 수 있는가?
  • RQ3이국적 데이터를 클러스터링 이전에 효과적으로 탐지하고 제거하여 강건성을 향상시킬 수 있는가?
  • RQ4LDPS를 $k$-medoids와 비유클리드 거리와 결합하면 비구형 또는 다양체 분포 데이터에서 성능을 향상시킬 수 있는가?
  • RQ5실세계 데이터셋에서 클러스터 수 추정 및 정확도 측면에서 최신 기술과 비교해 본 결과, 제안된 방법은 어떤가?

주요 결과

  • LDPS-means와 LDPS-medoids는 특히 진정한 클러스터 수 $k^*$가 클 경우 표준 $k$-means와 $k$-medoids보다 뛰어난 클러스터링 성능을 달성한다.
  • Olivetti Face Database(Oliv.-40)에서 LDPS-medoids는 이전 연구 대비 오류율 $r_e$를 15.9% 감소시키고, 거짓률 $r_f$를 25% 감소시켰다.
  • Oliv.-40에서 LDPS-medoids는 $r_t = 74.0\%$의 성능을 기록하여 기준 방법 대비 8.8% 향상되었다.
  • Oliv.-10, Oliv.-20, Oliv.-30에서 이 방법은 $x$-means, $dip$-means, CFSFDP보다 더 정확하게 $k$를 추정했으며, 진실값에 가깝게 일관되게 추정했다.
  • LDPS-medoids는 표준 $k$-means가 수천 번의 랜덤 재시작을 거쳐도 달성할 수 없는 수준의 SSE*를 더 적은 반복 횟수로 달성했다.
  • LDPS-medoids에서 고차원적이고 비구형인 데이터(예: 얼굴 이미지)에 대해 다양체 거리(기반으로 한 CW-SSIM과 $t$-nn 이웃)를 사용함으로써 성능 향상이著명하게 향상되었다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.