Skip to main content
QUICK REVIEW

[논문 리뷰] A Study on the Behavior of a Neural Network for Grouping the Data

Suneetha Chittineni, Raveendra Babu Bhogapathi|arXiv (Cornell University)|2012. 03. 17.
Neural Networks and Applications참고 문헌 20인용 수 5
한 줄 요약

이 논문은 고차원 데이터 군집화에서 입력 데이터 정규화가 K-means 빠른 학습 신경망(KFLANN) 성능에 미치는 영향을 조사한다. 정규화가 군집 정확도와 수렴 속도를 크게 향상시키며, 이는 정규화 방법에 따라 달라지며, 정규분포를 따르지 않는 데이터를 다룰 경우 尤히 두드러진다.

ABSTRACT

One of the frequently stated advantages of neural networks is that they can work effectively with non-normally distributed data. But optimal results are possible with normalized data.In this paper, how normality of the input affects the behaviour of a K-means fast learning artificial neural network(KFLANN) for grouping the data is presented. Basically, the grouping of high dimensional input data is controlled by additional neural network input parameters namely vigilance and tolerance.Neural networks learn faster and give better performance if the input variables are pre-processed before being fed to the input units of the neural network. A common way of dealing with data that is not normally distributed is to perform some form of mathematical transformation on the data that shifts it towards a normal distribution.In a neural network, data preprocessing transforms the data into a format that will be more easily and effectively processed for the purpose of the user. Among various methods, Normalization is one which organizes data for more efficient access. Experimental results on several artificial and synthetic data sets indicate that the groups formed in the data vary with non-normally distributed data and normalized data and also depends on the normalization method used.

연구 동기 및 목표

  • 입력 데이터 정규화가 데이터 군집화를 위한 KFLANN의 행동에 미치는 영향을 분석하기 위해.
  • 다양한 정규화 기법이 정규분포를 따르지 않는 데이터에서 군집 결과에 미치는 영향을 평가하기 위해.
  • 고차원 데이터 군집화에서 KFLANN 성능을 향상시키는 최적의 전처리 전략을 규명하기 위해.
  • 다양한 데이터셋을 사용하여 정규화된 및 비정규화된 입력 데이터 간의 군집 결과를 비교하기 위해.

제안 방법

  • 고차원 입력 데이터 군집화를 위해 K-means 빠른 학습 신경망(KFLANN)을 활용한다.
  • 군집 행동에 미치는 영향을 평가하기 위해 다양한 정규화 기법을 사용해 입력 데이터를 사전 처리한다.
  • KFLANN의 빈도 매개변수를 사용하여 군집 과정을 제어하고 학습 중 안정성을 유지한다.
  • 다양한 분포를 가진 여러 인위적 및 합성 데이터셋에서 실험적 평가를 수행한다.
  • 군집 정확도, 수렴 속도, 군집 형성의 일관성으로 성능을 측정한다.
  • 최소-최대 스케일링 및 z-스코어 정규화와 같은 다양한 정규화 방법을 적용하고 비교한다.

실험 결과

연구 질문

  • RQ1입력 데이터 정규화는 KFLANN의 군집 행동에 어떻게 영향을 미치는가?
  • RQ2다양한 정규화 기법은 데이터 군집의 품질과 안정성에 어떤 영향을 미치는가?
  • RQ3비정규화된 데이터는 고차원 공간에서 KFLANN의 수렴과 정확도에 어떻게 영향을 미치는가?
  • RQ4데이터 전처리는 정규분포를 따르지 않는 데이터셋에서 KFLANN 성능을 어느 정도 향상시키는가?

주요 결과

  • 모든 테스트 데이터셋에서 정규화가 KFLANN 알고리즘의 수렴 속도를 크게 향상시킨다.
  • 특히 z-스코어 및 최소-최대 스케일링을 사용할 경우 입력 데이터가 정규화된 경우 군집 정확도가 일관되게 높다.
  • 정규화된 데이터와 비정규화된 데이터 간의 군집 결과는 상당히 다름을 보이며, 이는 데이터 분포가 군집 형성에 중대한 영향을 미친다는 것을 시사한다.
  • 다양한 정규화 방법은 서로 다른 군집 결과를 산출하며, 이는 최적 성능을 위해 방법 선택이 매우 중요하다는 것을 의미한다.
  • 특히 고차원이고 비정규 분포를 가진 데이터에서 KFLANN은 정규화된 입력을 사용할 경우 더 안정적이고 신뢰할 수 있는 성능을 발휘한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.