Skip to main content
QUICK REVIEW

[論文レビュー] Enriching Complex Networks with Word Embeddings for Detecting Mild Cognitive Impairment from Speech Transcripts

Leandro Borges dos Santos, Edilson A. Corrêa|arXiv (Cornell University)|Jan 1, 2017
Topic Modeling参考文献 29被引用数 9
ひとこと要約

本稿では、スクリプトを複雑なネットワークとしてモデル化し、語の分散表現を統合することで意味的表現を向上させる、Mild Cognitive Impairment (MCI) の検出のための新規手法 Complex Networks Enriched with Word Embeddings (CNE) を提案する。DementiaBank および Cinderella の2つのデータセットにおいて、従来の Bag-of-Words および言語的特徴量アプローチを上回る高い精度を達成し、大規模な MCI スクリーニングへの応用可能性を示した。

ABSTRACT

Mild Cognitive Impairment (MCI) is a mental disorder difficult to diagnose.Linguistic features, mainly from parsers, have been used to detect MCI, but this is not suitable for large-scale assessments.MCI disfluencies produce nongrammatical speech that requires manual or high precision automatic correction of transcripts.In this paper, we modeled transcripts into complex networks and enriched them with word embedding (CNE) to better represent short texts produced in neuropsychological assessments.The network measurements were applied with well-known classifiers to automatically identify MCI in transcripts, in a binary classification task.A comparison was made with the performance of traditional approaches using Bag of Words (BoW) and linguistic features for three datasets: DementiaBank in English, and Cinderella and Arizona-Battery in Portuguese.Overall, CNE provided higher accuracy than using only complex networks, while Support Vector Machine was superior to other classifiers.CNE provided the highest accuracies for Dementia-Bank and Cinderella, but BoW was more efficient for the Arizona-Battery dataset probably owing to its short narratives.The approach using linguistic features yielded higher accuracy if the transcriptions of the Cinderella dataset were manually revised.Taken together, the results indicate that complex networks enriched with embedding is promising for detecting MCI in large-scale assessments.1 talkbank.org/DementiaBank/

研究の動機と目的

  • スクリプトからの大規模かつ自動的な Mild Cognitive Impairment (MCI) の検出の課題に対処すること。
  • 文法的でない、不順序な会話において意味のニュアンスを捉えることが難しい従来の言語的解析および Bag-of-Words (BoW) モデルの限界を克服すること。
  • 語の分散表現を複雑なネットワーク構造に統合することで、短い障害を呈する会話スクリプトの表現を改善すること。
  • CNE の性能を、複数の多言語データセットにおいて、標準的手法(BoW および言語的特徴量)と比較して評価すること。

提案手法

  • 語をノードとし、語の共起をエッジとする手法により、スクリプトを複雑なネットワークとしてモデル化した。
  • 語の分散表現(例:Word2Vec や類似手法)を用いて、ネットワークのノードに密な意味的表現を付加した。
  • ネットワーク測定値(例:次数中心性、クラスタ係数、媒介性)を特徴量ベクトルとして抽出した。
  • これらの特徴量を標準的な分類器に供給し、Support Vector Machine (SVM) が最も優れた性能を示した。
  • 評価は3つのデータセットで実施した:DementiaBank(英語)、Cinderella(ポルトガル語)、Arizona-Battery(ポルトガル語)。
  • BoW および言語的特徴量ベースラインと比較したが、Cinderella データセットについては言語的特徴量のための手動での修正を実施した。

実験結果

リサーチクエスチョン

  • RQ1語の分散表現を複雑なネットワークに統合することで、従来の手法と比較して、会話スクリプトにおける MCI の検出性能が向上するか?
  • RQ2CNE のアプローチは、異なる言語的および臨床的特徴を持つデータセットにおいて、Bag-of-Words や言語的特徴量ベースのモデルと比較してどのように性能を発揮するか?
  • RQ3スクリプトの手動による修正は、言語的特徴量ベースのモデルの MCI 検出性能を向上させるか?
  • RQ4CNE のアプローチは、スクリプト長や言語的特性が異なる多様なデータセットにおいても頑健か?

主な発見

  • CNE は、DementiaBank(英語)および Cinderella(ポルトガル語)データセットにおいて、BoW や言語的特徴量と比較して最高の精度を達成した。
  • 短いナラティブを含む Arizona-Battery データセットでは、BoW が CNE を上回った。これは、効果が文脈依存である可能性を示唆している。
  • テストされた分類器の中で、Support Vector Machine (SVM) が最も効果的であり、全データセットで他のモデルを上回った。
  • Cinderella データセットにおけるスクリプトの手動修正は、言語的特徴量ベースのモデルの精度を顕著に向上させた。
  • 語の分散表現を複雑なネットワークに統合することで、MCI に特徴的な不順序で文法的でない会話の意味的表現が向上した。
  • 全体として、CNE は臨床的および大規模な環境におけるスケーラブルで自動化された MCI 検出の強力な可能性を示している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。