Skip to main content
QUICK REVIEW

[논문 리뷰] Comment on "Comparing two formulations of skew distributions with special reference to model-based clustering" by A. Azzalini, R. Browne, M. Genton, and P. McNicholas

Geoffrey J. McLachlan, Sharon Lee|arXiv (Cornell University)|2014. 04. 07.
Statistical Distribution Estimation and Applications참고 문헌 9인용 수 3
한 줄 요약

이 논문은 모델 기반 군집화에서 비대칭 t-분포의 비교 연구를 비판하며, 참고된 논문에서 제시된 제한된 모형과 비제한 모형이 잘못 기술되어 있음을 주장한다. 실제 데이터셋에서 비제한 비대칭 t-혼합 모형이 보고된 값의 1/3 미만인 오분류율을 기록하며 훨씬 더 뛰어난 군집화 성능을 보임을 입증하고, 양 모형을 포함하는 더 유연하고 계산적으로 효율적인 대안으로 캐논리컬 기본 비대칭 t(CFUST) 분포 클래스를 제안한다.

ABSTRACT

In this paper, we comment on the recent comparison in Azzalini et al. (2014) of two different distributions proposed in the literature for the modelling of data that have asymmetric and possibly long-tailed clusters. They are referred to as the restricted and unrestricted skew t-distributions by Lee and McLachlan (2013a). Firstly, we wish to point out that in Lee and McLachlan (2014b), which preceded this comparison, it is shown how a distribution belonging to the broader class, the canonical fundamental skew t (CFUST) class, can be fitted with essentially no additional computational effort than for the unrestricted distribution. The CFUST class includes the restricted and unrestricted distributions as special cases. Thus the user now has the option of letting the data decide as to which model is appropriate for their particular dataset. Secondly, we wish to identify several statements in the comparison by Azzalini et al.(2014) that demonstrate a serious misunderstanding of the reporting of results in Lee and McLachlan (2014a) on the relative performance of these two skew t-distributions. In particular, there is an apparent misunderstanding of the nomenclature that has been adopted to distinguish between these two models. Thirdly, we take the opportunity to report here that we have obtained improved fits, in some cases a marked improvement, for the unrestricted model for various cases corresponding to different combinations of the variables in the two real datasets that were used in Azzalini et al. (2014) to mount their claims on the relative superiority of the restricted and unrestricted models. For one case the misclassification rate of our fit under the unrestricted model is less than one third of their reported error rate. Our results thus reverse their claims on the ranking of the restricted and unrestricted models in such cases.

연구 동기 및 목표

  • 아자린리 등(2014)이 수행한 제한형과 비제한형 비대칭 t-분포 비교에 대한 잘못된 기술을 수정하기 위해
  • 비제한 비대칭 t-혼합 모형이 참고된 연구에서 보고된 것보다 훨씬 뛰어난 군집화 성능을 보임을 입증하기 위해
  • 제한형과 비제한형 양 형태를 특수한 경우로 포함하는 통합적이고 계산적으로 효율적인 모델로서 캐논리컬 기본 비대칭 t(CFUST) 분포를 제안하기 위해
  • 비대칭 분포 연구에서 제한형과 비제한형 비대칭 t-분포의 차이를 명확히 하여 향후 연구에서 혼동을 방지하기 위해

제안 방법

  • 제한형과 비제한형 비대칭 t-분포를 특수한 경우로 포함하는 더 넓은 범주인 캐논리컬 기본 비대칭 t(CFUST) 분포를 제안한다.
  • CFUST 모형을 적합하는 데에는 비제한 모형과 거의 동일한 계산 비용이 소요됨을 입증한다.
  • 실제 데이터셋에 대해 CFUST 분포의 유한 혼합 모형(FM-CFUST)을 적용하여 모형 비교를 수행한다.
  • 비제한 모형을 사용하여 크랩 및 AIS 두 개의 실제 데이터셋을 재분석하여 아자린리 등(2014)에서 보고한 오분류율과 비교한다.
  • 모든 모형에 동일한 매개수 추정 프레임워크(EM 알고리즘)를 사용하여 공정한 비교를 확보한다.
  • 비제한 모형이 제한 모형에 포함되지 않으며, 따라서 성능 평가가 데이터셋 별로 이루어져야 하며 일반적으로 평가되어서는 안 된다는 점을 강조한다.

실험 결과

연구 질문

  • RQ1아자린리 등(2014)의 주장과는 반대로, 비제한 비대칭 t-혼합 모형이 실제 군집화 작업에서 제한 모형을 능가하는가?
  • RQ2비제한 모형과 비교해 볼 때 캐논리컬 기본 비대칭 t(CFUST) 분포를 적합하는 데 추가적인 계산 비용이 거의 들지 않는가?
  • RQ3아자린리 등(2014)에서 보고된 바에 따르면 비제한 모형의 성능가 잘못 기술된 이유는 무엇인가?
  • RQ4비대칭 분포 연구에서 명칭 혼동이 모형 성능 해석에 어떤 영향을 미치는가?
  • RQ5CFUST 모형이 제한형과 비제한형 모형 간의 데이터 기반 선택을 가능하게 하는 통합 프레임워크로 기능할 수 있는가?

주요 결과

  • 크랩 데이터셋에서 비제한 모형의 오분류율(MCR)은 0.11로, 아자린리 등(2014)에서 보고한 0.36의 1/3 미만이었으며, 이는 그들의 모형 우월성에 대한 결론을 뒤집는 결과였다.
  • AIS 데이터셋에서는 비제한 모형이 아자린리 등(2014)에서 보고한 것보다 약간 더 많은 이변량 및 삼변량 조합에서 제한 모형을 능가했으며, 이는 제한 모형이 광범위하게 우월하다는 그들의 주장에 반박되는 결과였다.
  • 비제한 모형은 참고된 연구에서 보고된 값보다 실제 데이터셋에서 훨씬 더 뛰어난 적합도를 보였으며, 이는 심각한 성능 격차가 있음을 시사한다.
  • CFUST 클래스를 통해 비제한 모형과 동일한 계산 비용으로 더 유연한 모형을 적합할 수 있으며, 이는 데이터 기반의 모형 선택을 가능하게 한다.
  • 리 등과 매클라클란(2013a)에서 사용된 명칭은 제한 모형이 단일변량 비대칭 함수를 갖는다는 점을 명확히 하며, 매개수 공간에 대한 제한이 아니라는 점을 반영한다. 이는 ABGM에서의 오해된 해석과는 다릅니다.
  • 저자들은 제한형 모형보다 비제한 모형이 크랩, AIS 등 여러 실제 데이터셋에서 일관되게 뛰어난 성능을 보임을 확인하였으며, 이 성능은 일반적인 경우로 일반화되지 않는다는 점을 확인한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.