Skip to main content
QUICK REVIEW

[論文レビュー] MFCC-based Recurrent Neural Network for Automatic Clinical Depression Recognition and Assessment from Speech

Emna Rejaibi, Ali Komaty|arXiv (Cornell University)|Sep 16, 2019
Emotion and Mood Recognition参考文献 38被引用数 246
ひとこと要約

MFCC ベースの LSTM フレームワークは音声からうつ病を検出し、PHQ-8 の重症度を推定します。データ拡張と転移学習を用いて小規模データセットの課題を克服し、うつ病検出の検証精度 76.27%、二項タスクの RMSE 0.405 を達成します。

ABSTRACT

Clinical depression or Major Depressive Disorder (MDD) is a common and serious medical illness. In this paper, a deep recurrent neural network-based framework is presented to detect depression and to predict its severity level from speech. Low-level and high-level audio features are extracted from audio recordings to predict the 24 scores of the Patient Health Questionnaire and the binary class of depression diagnosis. To overcome the problem of the small size of Speech Depression Recognition (SDR) datasets, expanding training labels and transferred features are considered. The proposed approach outperforms the state-of-art approaches on the DAIC-WOZ database with an overall accuracy of 76.27% and a root mean square error of 0.4 in assessing depression, while a root mean square error of 0.168 is achieved in predicting the depression severity levels. The proposed framework has several advantages (fastness, non-invasiveness, and non-intrusion), which makes it convenient for real-time applications. The performances of the proposed approach are evaluated under a multi-modal and a multi-features experiments. MFCC based high-level features hold relevant information related to depression. Yet, adding visual action units and different other acoustic features further boosts the classification results by 20% and 10% to reach an accuracy of 95.6% and 86%, respectively. Considering visual-facial modality needs to be carefully studied as it sparks patient privacy concerns while adding more acoustic features increases the computation time.

研究の動機と目的

  • MFCC特徴量とRNNを用いた音声からのうつ病検出と重症度評価を動機づける。
  • データ拡張と関連感情タスクからの転移学習により、小規模データセットの課題に対処する。
  • DAIC-WOZコーパスでの性能を評価し、性別・ノイズ耐性・一般化を分析する。

提案手法

  • 前処理済み音声セグメントから MFCC特徴量(60 次数)とその一階および二階微分を抽出する。
  • バイナリうつ病分類(シグモイド)と多クラスPHQ-8重症度推定(ソフトマックス)のために、3層 LSTM ネットワークの後に二つの全結合層を用いる。
  • 係数全体でグローバル z-score 正規化を用いて MFCC 特徴量を正規化する。
  • データ多様性を増すためノイズ、ピッチ、シフト、速度のデータ拡張を適用する。
  • 感情認識タスク(RAVDESS)で事前学習し、うつ病に対してファインチューニングする(転移学習)。
  • ベースライン、データ拡張、転移学習設定を比較し、二値うつ病と24レベルの重症度予測の両方を評価する。

実験結果

リサーチクエスチョン

  • RQ1MFCCベースのRNNは音声だけでうつ病を正確に検出できるか?
  • RQ2音声からPHQ-8の重症度レベルをどれくらい正確に予測できるか(二値 vs 多クラス)?
  • RQ3データ拡張と転移学習はDAIC-WOZでのうつ病検出と重症度推定を改善するか?
  • RQ4性別が音声からのうつ病認識におけるモデル性能にどのような影響を与えるか?
  • RQ5提案システムはノイズやデータセット間の一般化にどれくらい頑健か?

主な発見

  • ベースラインの MFCC ベース RNN はうつ病(二値)検出の検証精度 67.61%、RMSE 0.5057。
  • 重症度予測(24 PHQ-8 クラス)は RMSE 0.168、二値タスクより顕著に優れている。
  • データ拡張により検証精度が 74.0% に上昇し、RMSE が 0.4206(二値タスク)に低下。
  • RAVDESS の感情で事前学習し DAIC-WOZ でファインチューニングすると検証精度が 76.27%、RMSE が 0.4055(二値タスク)に向上。
  • 転移学習によりうつ病の F1 スコアが 38% から 46% に改善。性別分析では女性の方が検証精度が高い(85%)、男性は 83%。」
  • ノイズ耐性を示し、10% のガウスノイズで二値精度が 8.3% 減少し、20% ノイズまで安定した性能を示す。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。