Skip to main content
QUICK REVIEW

[논문 리뷰] Heart Disease Detection using Quantum Computing and Partitioned Random Forest Methods

Hanif Heidari, Gerhard Hellstern|arXiv (Cornell University)|2022. 08. 17.
Artificial Intelligence in Healthcare인용 수 5
한 줄 요약

이 논문은 2–4 큐비트를 사용하여 조기 심장병 진단을 향상시키기 위해 양자 컴퓨팅과 분할된 랜덤 포레스트를 통합한 하이브리드 양자 랜덤 포레스트(HQRF) 모델을 제안한다. 이 모델은 정확도와 강건성을 향상시킨다. Cleveland 데이터셋에서 AUC 점수 96.43%와 Statlog 데이터셋에서 97.78%를 기록하며, 이는 이전의 하이브리드 양자 신경망(HQNN)보다 작은 데이터셋과 큰 데이터셋 모두에서 외곽치에 대한 저항력과 효율성 면에서 뛰어나다.

ABSTRACT

Heart disease morbidity and mortality rates are increasing, which has a negative impact on public health and the global economy. Early detection of heart disease reduces the incidence of heart mortality and morbidity. Recent research has utilized quantum computing methods to predict heart disease with more than 5 qubits and are computationally intensive. Despite the higher number of qubits, earlier work reports a lower accuracy in predicting heart disease, have not considered the outlier effects, and requires more computation time and memory for heart disease prediction. To overcome these limitations, we propose hybrid random forest quantum neural network (HQRF) using a few qubits (two to four) and considered the effects of outlier in the dataset. Two open-source datasets, Cleveland and Statlog, are used in this study to apply quantum networks. The proposed algorithm has been applied on two open-source datasets and utilized two different types of testing strategies such as 10-fold cross validation and 70-30 train/test ratio. We compared the performance of our proposed methodology with our earlier algorithm called hybrid quantum neural network (HQNN) proposed in the literature for heart disease prediction. HQNN and HQRF outperform in 10-fold cross validation and 70/30 train/test split ratio, respectively. The results show that HQNN requires a large training dataset while HQRF is more appropriate for both large and small training dataset. According to the experimental results, the proposed HQRF is not sensitive to the outlier data compared to HQNN. Compared to earlier works, the proposed HQRF achieved a maximum area under the curve (AUC) of 96.43% and 97.78% in predicting heart diseases using Cleveland and Statlog datasets, respectively with HQNN. The proposed HQRF is highly efficient in detecting heart disease at an early stage and will speed up clinical diagnosis.

연구 동기 및 목표

  • 기존의 양자 기반 심장병 예측 모델의 한계를 해결하기 위해, 높은 큐비트 수요, 열악한 외곽치 처리 능력, 높은 계산 비용을 해결하고자 한다.
  • 작은 데이터셋과 큰 데이터셋 모두에 적합한 더 효율적이고 강건한 양자 기계학습 모델을 개발하고자 한다.
  • 양자 컴퓨팅과 분할된 랜덤 포레스트 접근 방식을 통합하여 예측 정확도와 AUC를 향상시키고자 한다.
  • 높은 성능을 유지하면서도 대규모 학습 데이터에 대한 의존도를 줄여 조기 임상 진단을 가능하게 하고자 한다.
  • 기존의 하이브리드 양자 신경망과 비교하여 모델의 외곽치에 대한 강건성(robustness)을 평가하고자 한다.

제안 방법

  • 제안된 HQRF 모델은 특징 학습과 분류 성능 향상을 위해 양자 회로와 분할된 랜덤 포레스트 앙상블을 결합한다.
  • 양자 회로는 2–4 큐비트로 구현되어 자원 소모를 최소화하면서도 높은 성능를 유지한다.
  • 데이터셋은 부분집합들로 분할되며, 각 부분집합은 랜덤 포레스트 프레임워크 내에서 양자 강화된 결정 트리에 의해 처리된다.
  • 랜덤 포레스트의 앙상블 특성 덕분에 외곽치 영향이 감소하여 극단적인 값에 대한 민감도가 낮아진다.
  • 두 가지 평가 전략을 사용한다: 10겹 교차 검증과 70-30 훈련/테스트 분할로, 성능 평가의 강건성을 확보한다.
  • 모델은 다양한 임상 데이터 분포를 반영하는 두 개의 오픈소스 데이터셋인 Cleveland와 Statlog에서 훈련 및 테스트된다.

실험 결과

연구 질문

  • RQ15 큐비트 이하의 양자 기계학습 모델이 기존의 양자 기반 모델보다 심장병 진단에서 더 높은 정확도를 달성할 수 있는가?
  • RQ2외곽치가 존재하는 상황에서 제안된 HQRF 모델은 HQNN 모델에 비해 어떻게 성능을 발휘하는가?
  • RQ3HQRF 모델은 작은 데이터셋과 큰 데이터셋 모두에서 높은 성능를 유지하는가?
  • RQ4분할된 랜덤 포레스트 아키텍처의 사용이 모델의 일반화 능력과 계산 효율성에 어떤 영향을 미치는가?
  • RQ5AUC 및 훈련 효율성 측면에서 HQRF 모델은 HQNN 모델에 비해 어떻게 비교되는가?

주요 결과

  • HQRF 모델은 Cleveland 데이터셋에서 AUC 96.43%를 기록하여 이전의 양자 모델을 능가했다.
  • Statlog 데이터셋에서는 AUC 97.78%를 기록하여 교차 데이터셋 평가에서 뛰어난 성능를 입증했다.
  • HQRF는 HQNN 모델보다 외곽치 데이터에 대한 민감도가著적으로 낮아 강건성이 향상되었다.
  • HQRF는 작은 데이터셋과 큰 데이터셋 모두에서 높은 성능를 유지했으며, 반면 HQNN은 최적 성능를 발휘하기 위해 대규모 데이터셋이 필요했다.
  • 단지 2–4 큐비트로도 높은 정확도를 달성하여 이전의 양자 모델에 비해 계산 비용과 메모리 사용량을 줄였다.
  • 10겹 교차 검증에서는 HQRF가 HQNN을 앞섰지만, 70-30 훈련/테스트 분할에서는 HQNN이 더 우수한 성능를 보여, 데이터셋 크기 의존성의 존재를 시사했다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.