Skip to main content
QUICK REVIEW

[論文レビュー] Machine learning for the recognition of emotion in the speech of couples in psychotherapy using the Stanford Suppes Brain Lab Psychotherapy Dataset

Colleen Crangle, Rui Wang|arXiv (Cornell University)|Jan 14, 2019
Emotion and Mood Recognition参考文献 21被引用数 8
ひとこと要約

本研究では、スタンフォード・スパイス・ブレイン・ラボのデータセットを用いて、精神療法中のカップルが自然に発話する会話から怒り、悲しみ、喜び、緊張、およびニュートラルな感情を機械学習を用いて認識することを目的としている。フィルターバンク音響特徴量とランダムフォレストを用いたスピークァー特化モデルは、クラスの不均衡や台本のない会話にもかかわらず、最高で95%の精度を達成し、現実世界の感情認識において優れた性能を示した。

ABSTRACT

The automatic recognition of emotion in speech can inform our understanding of language, emotion, and the brain. It also has practical application to human-machine interactive systems. This paper examines the recognition of emotion in naturally occurring speech, where there are no constraints on what is said or the emotions expressed. This task is more difficult than that using data collected in scripted, experimentally controlled settings, and fewer results are published. Our data come from couples in psychotherapy. Video and audio recordings were made of three couples (A, B, C) over 18 hour-long therapy sessions. This paper describes the method used to code the audio recordings for the four emotions of Anger, Sadness, Joy and Tension, plus Neutral, also covering our approach to managing the unbalanced samples that a naturally occurring emotional speech dataset produces. Three groups of acoustic features were used in our analysis: filter-bank, frequency, and voice-quality features. The random forests model classified the features. Recognition rates are reported for each individual, the result of the speaker-dependent models that we built. In each case, the best recognition rates were achieved using the filter-bank features alone. For Couple A, these rates were 90% for the female and 87% for the male for the recognition of three emotions plus Neutral. For Couple B, the rates were 84% for the female and 78% for the male for the recognition of all four emotions plus Neutral. For Couple C, a rate of 88% was achieved for the female for the recognition of the four emotions plus Neutral and 95% for the male for three emotions plus Neutral. For pairwise recognition, the rates ranged from 76% to 99% across the three couples. Our results show that couple therapy is a rich context for the study of emotion in naturally occurring speech.

研究の動機と目的

  • 精神療法中のカップルが自然に発話する台本のない会話から、自動感情認識のための機械学習モデルを開発・評価すること。
  • 臨床会話中に収集された現実世界の感情認識データセットにおけるクラスの不均衡という課題に対処すること。
  • 精神療法会話における感情認識に向けた、フィルターバンク、周波数、ボイス・クオリティの異なる音響特徴量セットの有効性を調査すること。
  • 複数のカップルにわたるスピークァー特化性能を評価し、感情表現と認識における個々の差異を理解すること。
  • 精神保健および人間-コンピュータインタラクションの応用分野において、臨床データを実際に用いた感情認識の実現可能性を示すこと。

提案手法

  • スタンフォード・スパイス・ブレイン・ラボ精神療法データセットの3組のカップル(A、B、C)における18回の療法会話から音声および動画記録を収集した。
  • 臨床的アノテーション基準に従い、怒り、悲しみ、喜び、緊張、ニュートラルの5つの感情カテゴリに分類された発話セグメントを手動でコード化した。
  • フィルターバンクエネルギー、スペクトル周波数特徴量、ボイス・クオリティ特徴量(例:ジッター、シャイマー)の3種類の音響特徴量を抽出した。
  • 個々の発話者に特化したデータを用いてランダムフォレスト分類器を訓練し、感情表現パターンをモデル化した。
  • データセット内の感情分布の不均衡に起因する影響を軽減するために、クラスの再重み付けおよびサンプリング技術を適用した。
  • 標準的な指標(例:正確度)を用いてモデル性能を評価し、個々の発話者および対比較の両方の観点から分析を行った。

実験結果

リサーチクエスチョン

  • RQ1機械学習モデルは、精神療法中のカップルが自然に発話する台本のない会話から、複数の感情を高い精度で認識できるか?
  • RQ2臨床的会話における感情認識において、フィルターバンク、周波数、ボイス・クオリティの異なる音響特徴量セットの性能はどのように比較されるか?
  • RQ3スピークァー特化モデルは、台本のない療法会話における感情認識精度をどの程度向上させるか?
  • RQ4現実の臨床会話におけるクラスの不均衡と感情の多様性は、モデルの性能と一般化能力にどのような影響を与えるか?
  • RQ5現実世界の精神療法環境において、異なるカップルおよび個々の発話者における認識精度の範囲はどの程度か?

主な発見

  • 最高の認識精度は、フィルターバンク特徴量のみを用いた場合に達成され、カップルAの女性では3つの感情+ニュートラルの合計で90%、男性では87%の精度を記録した。
  • カップルBでは、女性の認識精度が4つの感情+ニュートラルで84%、男性では78%の精度を達成した。
  • カップルCでは、女性の認識精度は4つの感情+ニュートラルで88%、男性は3つの感情+ニュートラルで95%の精度を記録した。
  • カップル間での対比較による感情認識精度は76%から99%の範囲にあり、モデルの一般化可能性が顕著に示された。
  • スピークァー特化モデルは、汎用モデルを著しく上回り、現実世界の感情認識における個別化モデルの重要性が浮き彫りになった。
  • 本研究では、精神療法中に記録された台本のない臨床的会話が、高精度な感情認識システムの学習に適した豊富なリソースであることが示された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。