Skip to main content
QUICK REVIEW

[논문 리뷰] Extracting the Locus of Attention at a Cocktail Party from Single-Trial EEG using a Joint CNN-LSTM Model.

Ivine Kuruvila, Jan Muncke|arXiv (Cornell University)|2021. 02. 08.
EEG and Brain-Computer Interfaces참고 문헌 28인용 수 6
한 줄 요약

이 논문은 단일 시행 EEG 신호와 다중화자 스펙트로그램을 분석하여 코ctail party 환경에서 听覚 주의를 추론하는 공동 CNN-LSTM 모델을 제안한다. 3초 간의 중앙값 복원 정확도는 77.2%를 기록한다. 모델은 희박성에 대해 강건성을 보이며, 크기 기반 50% 프루닝 조건에서도 성능을 유지한다.

ABSTRACT

Human brain performs remarkably well in segregating a particular speaker from interfering speakers in a multi-speaker scenario. It has been recently shown that we can quantitatively evaluate the segregation capability by modelling the relationship between the speech signals present in an auditory scene and the cortical signals of the listener measured using electroencephalography (EEG). This has opened up avenues to integrate neuro-feedback into hearing aids whereby the device can infer user's attention and enhance the attended speaker. Commonly used algorithms to infer the auditory attention are based on linear systems theory where the speech cues such as envelopes are mapped on to the EEG signals. Here, we present a joint convolutional neural network (CNN) - long short-term memory (LSTM) model to infer the auditory attention. Our joint CNN-LSTM model takes the EEG signals and the spectrogram of the multiple speakers as inputs and classifies the attention to one of the speakers. We evaluated the reliability of our neural network using three different datasets comprising of 61 subjects where, each subject undertook a dual-speaker experiment. The three datasets analysed corresponded to speech stimuli presented in three different languages namely German, Danish and Dutch. Using the proposed joint CNN-LSTM model, we obtained a median decoding accuracy of 77.2% at a trial duration of three seconds. Furthermore, we evaluated the amount of sparsity that our model can tolerate by means of magnitude pruning and found that the model can tolerate up to 50% sparsity without substantial loss of decoding accuracy.

연구 동기 및 목표

  • 다중화자 환경에서 EEG 신호로부터 听覚 주의를 추론하는 딥러닝 모델을 개발한다.
  • 원시 EEG 및 스펙트로그램 입력으로부터 엔드 투 엔드 학습을 활용하여 선형 모델을 향상시킨다.
  • 다양한 데이터셋을 사용하여 언어 간 일반화 능력을 평가한다.
  • 크기 기반 프루닝을 통한 구조적 희박성에 대한 모델 강건성 평가

제안 방법

  • 공동 합성곱 신경망(CNN)과 장기 단기 기억(LSTM) 아키텍처가 다중 화자들의 EEG 및 스펙트로그램을 처리한다.
  • CNN은 EEG에서 공간적 및 시간적 특징을 추출하고, LSTM은 신경 반응의 순차적 의존성을 모델링한다.
  • 입력 특징으로는 독일어, 덴마크어, 네덜란드어로 된 이중화자 음성 자극의 EEG 시계열과 스펙트로그램이 포함된다.
  • 통합 입력 표현을 기반으로 청취자가 주의를 기울이고 있는 화자를 분류한다.
  • 정확도를 모니터링하면서 매개변수 수를 줄이며 크기 기반 프루닝을 통해 모델 강건성을 평가한다.

실험 결과

연구 질문

  • RQ1공동 CNN-LSTM 모델은 다중화자 상황에서 단일 시행 EEG로부터 정확하게 听覚 주의를 복원할 수 있는가?
  • RQ2이중화자 실험에서 모델의 성능은 다양한 언어 간에 어떻게 달라지는가?
  • RQ3모델은 정확도 손실이 심각하게 발생하지 않는 수준까지 얼마나 많은 구조적 희박성에 견딜 수 있는가?
  • RQ4공동 아키텍처는 전통적인 선형 모델보다 주의 복원에서 더 우수한 성능을 보이는가?

주요 결과

  • 공동 CNN-LSTM 모델은 61명의 피실험자 대상으로 3초 간의 시행 기간 동안 중앙값 복원 정확도 77.2%를 달성했다.
  • 독일어, 덴마크어, 네덜란드어 등 세 가지 다른 언어 데이터셋에서 높은 성능 유지를 보이며, 언어 간 일반화 능력을 입증했다.
  • 크기 기반 프루닝 분석을 통해 모델이 정확도 저하 없이 최대 50%의 희박성에 견딜 수 있음을 입증했다.
  • 결과적으로 딥러닝 모델이 단일 시행 EEG로부터 听覚 주의를 효과적으로 복원할 수 있으며, 실시간 뇌신호 피드백 응용이 가능하다는 점을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.