[논문 리뷰] Machine learning for the recognition of emotion in the speech of couples in psychotherapy using the Stanford Suppes Brain Lab Psychotherapy Dataset
이 연구는 스탠퍼드 수퍼스 브레인 랩 데이터셋을 사용하여 심리치료 상황에서 부부 간 자연스러운 대화에서 분노, 슬픔, 기쁨, 긴장, 중립 등의 감정을 인식하기 위해 머신러닝을 적용한다. 필터백 음향 특징과 랜덤 포레스트를 사용한 발화자 독립 모델이 최대 95%의 정확도를 기록하여, 클래스 불균형과 즉석 대화의 영향을 받는 실제 환경에서의 감정 인식에 높은 성능을 보였다.
The automatic recognition of emotion in speech can inform our understanding of language, emotion, and the brain. It also has practical application to human-machine interactive systems. This paper examines the recognition of emotion in naturally occurring speech, where there are no constraints on what is said or the emotions expressed. This task is more difficult than that using data collected in scripted, experimentally controlled settings, and fewer results are published. Our data come from couples in psychotherapy. Video and audio recordings were made of three couples (A, B, C) over 18 hour-long therapy sessions. This paper describes the method used to code the audio recordings for the four emotions of Anger, Sadness, Joy and Tension, plus Neutral, also covering our approach to managing the unbalanced samples that a naturally occurring emotional speech dataset produces. Three groups of acoustic features were used in our analysis: filter-bank, frequency, and voice-quality features. The random forests model classified the features. Recognition rates are reported for each individual, the result of the speaker-dependent models that we built. In each case, the best recognition rates were achieved using the filter-bank features alone. For Couple A, these rates were 90% for the female and 87% for the male for the recognition of three emotions plus Neutral. For Couple B, the rates were 84% for the female and 78% for the male for the recognition of all four emotions plus Neutral. For Couple C, a rate of 88% was achieved for the female for the recognition of the four emotions plus Neutral and 95% for the male for three emotions plus Neutral. For pairwise recognition, the rates ranged from 76% to 99% across the three couples. Our results show that couple therapy is a rich context for the study of emotion in naturally occurring speech.
연구 동기 및 목표
- 자연스럽고 즉석의 대화에서 부부의 심리치료 상황에서의 자동 감정 인식을 위한 머신러닝 모델을 개발하고 평가하는 것.
- 임상 세션 기간에 수집된 실제 감정 음성 데이터셋에서의 클래스 불균형 문제를 해결하는 것.
- 치료적 대화에서 감정 인식에 효과적인 다양한 음향 특징 세트(필터백, 주파수, 음질)의 성능을 조사하는 것.
- 다양한 부부 간 발화자 독립 성능을 평가하여 감정 표현과 인식의 개인적 변동성을 이해하는 것.
- 정신건강 및 인간-컴퓨터 상호작용 응용 분야에서 실제 임상 데이터를 감정 인식에 활용할 수 있는 가능성을 입증하는 것.
제안 방법
- 스탠퍼드 수퍼스 브레인 랩 심리치료 데이터셋에서 3명의 부부(A, B, C)의 18회의 치료 세션에서 음성 및 영상 기록을 확보하였다.
- 임상 애너테이션 기준에 따라 분노, 슬픔, 기쁨, 긴장, 중립의 5개 감정 카테고리로 음성 세그먼트를 수동으로 애너테이션하였다.
- 필터백 에너지, 스펙트럼 주파수 특징, 음질 특징(예: 지터, 샤이머)의 3종류의 음향 특징을 추출하였다.
- 개별 발화자의 감정 표현 패턴을 모델링하기 위해 발화자별 데이터 기반 랜덤 포레스트 분류기를 훈련시켰다.
- 데이터셋 내 감정 분포의 불균형 영향을 줄이기 위해 클래스 재가중 및 샘플링 기법을 적용하였다.
- 정확도와 같은 표준 지표를 사용하여 성능을 평가하였으며, 개별 발화자 및 쌍별 비교 분석을 별도로 실시하였다.
실험 결과
연구 질문
- RQ1머신러닝 모델은 심리치료 상황에서 자연스럽고 즉석의 대화에서 여러 감정을 높은 정확도로 인식할 수 있는가?
- RQ2임상적 대화에서 감정 인식에 있어 필터백, 주파수, 음질 특징 세트의 성능는 어떻게 비교되는가?
- RQ3발화자 독립 모델링이 즉석 치료 대화에서 감정 인식 정확도를 얼마나 향상시키는가?
- RQ4실제 임상 세션에서의 클래스 불균형과 감정의 변동성은 모델 성능과 일반화 능력에 어떤 영향을 미치는가?
- RQ5실제 심리치료 환경에서 다양한 부부와 개인 발화자 간의 감정 인식 정확도 범위는 무엇인가?
주요 결과
- 가장 높은 인식 정확도는 필터백 특징만을 사용했을 때 달성되었으며, 부부 A의 여성은 3개 감정과 중립을 포함해 90%의 정확도를 기록했고, 남성은 87%였다.
- 부부 B의 경우 여성은 4개 감정과 중립을 포함해 84%의 정확도를 기록했고, 남성은 78%였다.
- 부부 C에서는 여성은 4개 감정과 중립을 포함해 88%의 정확도를 기록했고, 남성은 3개 감정과 중립을 포함해 95%의 정확도를 기록했다.
- 부부 간 쌍별 감정 인식 정확도는 76%에서 99%까지 다양했으며, 이는 개인 간 일반화 잠재력이 높다는 것을 시사한다.
- 발화자 독립 모델은 일반 모델보다 유의미하게 뛰어난 성능을 보였으며, 실제 감정 음성 인식에서 개인 맞춤형 모델링의 중요성을 입증하였다.
- 연구는 치료 상황에서 기록된 즉석의 임상 음성 자료가 고정확도 감정 인식 시스템을 훈련시키기에 실현 가능하고 풍부한 자료임을 입증하였다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.