Skip to main content
QUICK REVIEW

[論文レビュー] Learning Disease vs Participant Signatures: a permutation test approach to detect identity confounding in machine learning diagnostic applications

Elias Chaibub Neto, Abhishek Pratap|arXiv (Cornell University)|Dec 8, 2017
Imbalanced Data Classification Techniques参考文献 9被引用数 7
ひとこと要約

本稿では、機械学習診断アプリケーションにおけるアイデンティティ混同(identity confounding)を検出するための順列検定(permutation tests)を提案する。ここで、分類器が疾患パターンではなく被験者固有の特徴を学習してしまう可能性がある。記録単位の分割と被験者単位の分割における性能を比較することで、予測精度が疾患認識ではなく被験者特定に起因しているかどうかを同定する。結果として、音声データではアイデンティティ混同の強い証拠が得られ、被験者単位の分割をベストプラクティスとして支持する。

ABSTRACT

Recently, Saeb et al (2017) showed that, in diagnostic machine learning applications, having data of each subject randomly assigned to both training and test sets (record-wise data split) can lead to massive underestimation of the cross-validation prediction error, due to the presence of "subject identity confounding" caused by the classifier's ability to identify subjects, instead of recognizing disease. To solve this problem, the authors recommended the random assignment of the data of each subject to either the training or the test set (subject-wise data split). The adoption of subject-wise split has been criticized in Little et al (2017), on the basis that it can violate assumptions required by cross-validation to consistently estimate generalization error. In particular, adopting subject-wise splitting in heterogeneous data-sets might lead to model under-fitting and larger classification errors. Hence, Little et al argue that perhaps the overestimation of prediction errors with subject-wise cross-validation, rather than underestimation with record-wise cross-validation, is the reason for the discrepancies between prediction error estimates generated by the two splitting strategies. In order to shed light on this controversy, we focus on simpler classification performance metrics and develop permutation tests that can detect identity confounding. By focusing on permutation tests, we are able to evaluate the merits of record-wise and subject-wise data splits under more general statistical dependencies and distributional structures of the data, including situations where cross-validation breaks down. We illustrate the application of our tests using synthetic and real data from a Parkinson's disease study.

研究の動機と目的

  • 機械学習を用いた疾患診断における記録単位の分割と被験者単位の分割の対立について解決すること。
  • 分類器の性能が疾患状態ではなく被験者IDに起因しているかどうかを検出できる統計的手法を開発すること。
  • 記録単位と被験者単位の交差検証結果の乖離が、アイデンティティ混同によるものか、またはモデルの不足適合(under-fitting)によるものかを評価すること。
  • デジタルヘルスアプリケーションにおけるデータ分割戦略の妥当性を検証する実践的で実証的根拠に基づくアプローチを提供すること。

提案手法

  • 分類器の性能が被験者IDに依存しているかどうかを評価するための順列検定を提案する。
  • 被験者単位と記録単位のデータ分割を用いて、異なる仮定下での分類性能を比較する。
  • 疾患ラベルをシャッフルしながら被験者IDを保持する順列検定を適用し、性能が著しく低下するかどうかを評価する。
  • AUCと正答率(accuracy)を順列検定における性能指標として用い、アイデンティティ混同を検出する。
  • 合成データおよび実際のmPower研究データ(音声およびタッピング特徴)にこの手法を適用し、結果の妥当性を検証する。
  • 被験者単位の分割がアイデンティティ混同を軽減し、一般化性能を向上させることを示す。特に被験者数が十分に多い場合に顕著である。

実験結果

リサーチクエスチョン

  • RQ1記録単位のデータ分割における分類器の性能は、真の疾患認識を反映しているのか、それとも被験者IDによる混同に起因しているのか?
  • RQ2縦断的デジタルヘルスデータを用いて訓練された機械学習モデルにおいて、順列検定はアイデンティティ混同を信頼性高く検出できるか?
  • RQ3記録単位と被験者単位の交差検証結果の乖離は、アイデンティティ混同によるものか、それともモデルの不足適合によるものか?
  • RQ4被験者数が被験者単位の分割下での分類器の一般化能力にどのように影響するか?
  • RQ5特定のデータモダリティ(例:音声対タッピング)は、他のものよりもアイデンティティ混同にさらされやすいか?

主な発見

  • 順列検定により、記録単位の分割で訓練された音声ベースの分類器において、疾患ラベルをシャッフルしてもアイデンティティ混同の強い証拠が示された。
  • 記録単位の分割下でAUC値が高くなるのは、訓練データに100人以上の被験者が含まれている場合に限られ、疾患シグナルの検出が遅延していることを示唆している。
  • タッピングベースの分類器はアイデンティティ混同が少なく、被験者単位の分割下でより良好な一般化性能を示しており、より頑健な特徴表現であると考えられる。
  • 被験者単位の分割下でモデルの不足適合が観察されたが、訓練データに被験者数を増やすことで緩和可能であった。
  • 本研究では、性能の差がモデルの不足適合によるものであるという仮説に支持は得られなかった。
  • 音声特徴は、背景ノイズやマイクとの距離といった生物学的でないアーティファクトに感受しやすく、アイデンティティ混同にさらされやすいことが示唆された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。