Skip to main content
QUICK REVIEW

[논문 리뷰] Heart Sound Classification Considering Additive Noise and Convolutional Distortion

Farhat Binte Azam, Md. Istiaq Ansari|arXiv (Cornell University)|2021. 06. 03.
Phonocardiography and Auscultation Techniques인용 수 7
한 줄 요약

이 논문은 추가 노이즈와 센서 유도 콘볼루션 왜곡 하에서 강건한 심장음 분류를 위한 선형 필터뱅크(Fbank)와 멜 주파수 체프스트럼 계수(MFCC) 특징의 새로운 융합 방법을 제안한다. Fbank와 MFCC-13 특징을 융합한 잔차 컨볼루션 신경망(ResNet)을 사용하여, AUC 91.36%, F1 스코어 84.09%, Macc 85.08%를 달성하였으며, 이는 기존 방법들보다 뚜렷이 뛰어나며 특히 노이즈가 많은 환경과 다수의 stethoscope 환경에서 성능이 뛰어나다.

ABSTRACT

Cardiac auscultation is an essential point-of-care method used for the early diagnosis of heart diseases. Automatic analysis of heart sounds for abnormality detection is faced with the challenges of additive noise and sensor-dependent degradation. This paper aims to develop methods to address the cardiac abnormality detection problem when both types of distortions are present in the cardiac auscultation sound. We first mathematically analyze the effect of additive and convolutional noise on short-term filterbank-based features and a Convolutional Neural Network (CNN) layer. Based on the analysis, we propose a combination of linear and logarithmic spectrogram-image features. These 2D features are provided as input to a residual CNN network (ResNet) for heart sound abnormality detection. Experimental validation is performed on an open-access heart sound abnormality detection dataset involving noisy recordings obtained from multiple stethoscope sensors. The proposed method achieves significantly improved results compared to the conventional approaches, with an area under the ROC (receiver operating characteristics) curve (AUC) of 91.36%, F-1 score of 84.09%, and Macc (mean of sensitivity and specificity) of 85.08%. We also show that the proposed method shows the best mean accuracy across different source domains including stethoscope and noise variability, demonstrating its effectiveness in different recording conditions. The proposed combination of linear and logarithmic features along with the ResNet classifier effectively minimizes the impact of background noise and sensor variability for classifying phonocardiogram (PCG) signals. The proposed method paves the way towards developing computer-aided cardiac auscultation systems in noisy environments using low-cost stethoscopes.

연구 동기 및 목표

  • 실제 포인트온케어 환경에서 성능을 저하시키는 추가 노이즈와 콘볼루션 왜곡의 병합된 과제를 해결한다.
  • 다양한 stethoscope 센서와 노이즈 있는 기록 조건에서 자동 심장음 분석의 강건성을 향상시킨다.
  • 배경 노이즈와 센서 변동성의 영향을 최소화하는 특징 공학 접근법을 개발한다.
  • 기존 방법들과 비교해 소스 도메인(다양한 stethoscope와 노이즈 수준)에서 더 뛰어난 일반화 성능을 보여준다.
  • 딥 러닝 프레임워크 내에서 선형 및 로그 주파수 스펙트로그램 특징의 융합이 PCG 신호 분류에 효과적인지 검증한다.

제안 방법

  • 추가 왜곡과 콘볼루션 왜곡이 단기 필터뱅크 특징과 CNN 레이어에 미치는 영향을 수학적으로 분석하여 강건한 특징 설계에 기여한다.
  • PCG 신호에서 선형 필터뱅크 에너지(Fbank)와 멜 주파수 체프스트럼 계수(MFCC-13) 특징을 추출하여 상호 보완적인 표현을 확보한다.
  • Fbank와 MFCC-13 특징을 2차원 입력 행렬로 융합하여 딥 러닝을 위한 시간적 및 주파수적 동적 특성을 유지한다.
  • 융합된 스펙트로그램 특징에서 계층적 패턴을 학습하기 위해 잔차 컨볼루션 신경망(ResNet)을 분류기로 활용한다.
  • 다양한 stethoscope 센서와 노이즈 조건 간 공정한 평가를 보장하기 위해 도메인 균형 트레이닝(DBT)을 적용한다.
  • 기존 기준 모델 대비 성능 향상 여부를 통계적으로 검증하기 위해 McNemar의 카이제곱 검정을 적용한다.

실험 결과

연구 질문

  • RQ1추가 노이즈와 콘볼루션 왜곡이 심장음 분류 시스템의 성능에 함께 미치는 영향은 무엇인가?
  • RQ2선형 필터뱅크 에너지와 로그 멜 스펙트로그램 특징의 융합이 노이즈와 센서 변동성에 대한 강건성을 향상시킬 수 있는가?
  • RQ3제안된 특징 융합 전략이 노이즈가 많은 환경과 다중 소스 환경에서 기존 방법들보다 통계적으로 유의미한 성능 향상을 이끌 수 있는가?
  • RQ4기존 접근법들과 비교해 제안된 방법이 다양한 stethoscope 센서와 노이즈 수준에서 일반화 성능은 어떻게 되는가?
  • RQ5분류 성능 향상은 무작위 변동성 또는 데이터 泄漏 때문이 아니라 특징 융합 전략 탓인가?

주요 결과

  • 제안된 방법은 수신자 작동 특성 곡선 아래 면적(AUC)이 91.36%로 기준 모델 대비 뚜렷한 향상을 보였다.
  • 이 방법은 F1 스코어 84.09%와 민감도 및 특이도의 평균(Macc) 85.08%를 달성하여 강력한 종합 성능을 보였다.
  • Fbank와 MFCC-13 특징 융합은 McNemar 검정을 통해 기준 시스템(Potes-CNN-DBT) 대비 통계적으로 유의미한 향상(p < 0.05)을 보였다.
  • 다양한 소스 도메인(다양한 stethoscope와 노이즈 수준 포함)에서 평균 정확도가 가장 높아 일반화 능력이 뛰어나다는 것을 보여주었다.
  • Humayun 등(2023)의 tConv-CNN-DBT 시스템 대비 Macc에서 3.59% 향상되었지만 통계적으로 유의미하지 않음(p > 0.05)으로, 일관된 성능 향상이지만 극단적인 향상은 아니었다.
  • TSNE 시각화 결과 소스 도메인 별로 의미 있는 클러스터링이 관찰되지 않아, 도메인 이동에도 불구하고 모델이 불변 표현을 학습하고 있음을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.