Skip to main content
QUICK REVIEW

[論文レビュー] Unsupervised Machine Learning for the Discovery of Latent Disease Clusters and Patient Subgroups Using Electronic Health Records

Yanshan Wang, Yiqing Zhao|arXiv (Cornell University)|May 17, 2019
Machine Learning in Healthcare参考文献 32被引用数 6
ひとこと要約

本論文は、電子的健康記録(EHR)から潜在的な疾患クラスターや患者サブグループを、教師なし機械学習を用いて同定するための新しいポissonディリクレモデル(PDM)を提案する。標準的なLDAとは異なり、PDMは診断をポアソン分布でモデル化し、年齢と性別による期待診断数と観察された診断数を比較することで、人口統計的バイアスを低減しながら生物学的に意味のある共存疾患パターンを明らかにする。

ABSTRACT

Machine learning has become ubiquitous and a key technology on mining electronic health records (EHRs) for facilitating clinical research and practice. Unsupervised machine learning, as opposed to supervised learning, has shown promise in identifying novel patterns and relations from EHRs without using human created labels. In this paper, we investigate the application of unsupervised machine learning models in discovering latent disease clusters and patient subgroups based on EHRs. We utilized Latent Dirichlet Allocation (LDA), a generative probabilistic model, and proposed a novel model named Poisson Dirichlet Model (PDM), which extends the LDA approach using a Poisson distribution to model patients' disease diagnoses and to alleviate age and sex factors by considering both observed and expected observations. In the empirical experiments, we evaluated LDA and PDM on three patient cohorts with EHR data retrieved from the Rochester Epidemiology Project (REP), for the discovery of latent disease clusters and patient subgroups. We compared the effectiveness of LDA and PDM in identifying latent disease clusters through the visualization of disease representations learned by two approaches. We also tested the performance of LDA and PDM in differentiating patient subgroups through survival analysis, as well as statistical analysis. The experimental results show that the proposed PDM could effectively identify distinguished disease clusters by alleviating the impact of age and sex, and that LDA could stratify patients into more differentiable subgroups than PDM in terms of p-values. However, the subgroups discovered by PDM might imply the underlying patterns of diseases of greater interest in epidemiology research due to the alleviation of age and sex. Both unsupervised machine learning approaches could be leveraged to discover patient subgroups using EHRs but with different foci.

研究の動機と目的

  • ラベルなしの結果に依存せずにEHRデータ内の潜在的疾患クラスターや患者サブグループを同定すること。
  • 高齢化集団で一般的に観察される共存疾患パターンに影響を及ぼす年齢と性別の交絡要因を解消すること。
  • 期待される疾患数を組み込むことでLDAを拡張した、新しい教師なしモデル、ポアソンディリクレモデル(PDM)の開発と評価すること。
  • 実世界のEHRデータを用いて、PDMとLDAの両者が生物学的に意味のある疾患クラスターや患者サブグループを同定する性能を比較すること。
  • 生存分析と共存疾患プロファイルを用いて、同定されたサブグループの臨床的および疫学的妥当性を評価すること。

提案手法

  • 疾患クラスタを診断の確率的分布として表現する教師なしトピックモデルとしての、適応版ラティントディリクレ(LDA)をベースラインとして採用。
  • 診断データのカウントベースの性質を扱うために、ポアソン分布を用いて患者の診断をモデル化する、新しいポアソンディリクレモデル(PDM)を提案。
  • 年齢・性別グループごとの観察済みおよび期待される疾患数を組み込み、疾患有病率における人口統計的バイアスを補正。
  • 潜在的疾患クラスタの確率的推論を可能にするために、疾患トピック分布にディリクレ事前分布を適用。
  • LDAおよびPDMの両モデルにおけるパラメータ推定に、マルコフ連鎖モンテカルロ(MCMC)手法を適用。
  • LDAとPDMの間でのクラスタ分離度を比較するため、2次元の潜在トピック空間に疾患表現を可視化。

実験結果

リサーチクエスチョン

  • RQ1提案されたPDMは、年齢と性別の交絡効果を最小限に抑えつつ、EHRデータ内に明確な潜在的疾患クラスタを効果的に同定できるか?
  • RQ2PDMが同定する患者サブグループは、LDAが同定するものと比較して、臨床的および人口統計的特徴でどれほど区別可能か?
  • RQ3PDMが同定する患者サブグループは、年齢と性別を超えて生物学的または疫学的に意味のあるパターンを反映しているか?
  • RQ4LDAとPDMは生存分析において、予後分類の観点からどれほど患者を効果的に分類できるか?
  • RQ5各モデルが同定するサブグループ間で、エリクサウアー共存疾患インデックス(ECI)スコアにどれほど差が生じるか?

主な発見

  • PDMは、年齢と性別に基づく期待疾患有病率を考慮することで、LDAよりも明確で生物学的に解釈可能な疾患クラスタを同定した。
  • 生存分析において、LDAはPDMよりも患者サブグループをより効果的に分類し、p < 0.001の低いp値を示した。これは統計的差別化が強いことを示している。
  • 生存分析における統計的有意性がやや低かったにもかかわらず、PDMが導いたサブグループは、年齢と性別の交絡要因を低減しており、疫学的により意味のある関連性を示した。
  • PDMが同定したサブグループ間でエリクサウアー共存疾患インデックス(ECI)スコアに有意差が認められ、共存疾患負荷における臨床的差異が示された。
  • 疾患表現の可視化により、PDMはLDAよりも潜在トピック空間でより明確で、より二分化されたクラスタを生成した。
  • 提案されたPDMモデルは、年齢と性別が主な交絡要因となる高齢者集団において、潜在的疾患パターンを同定するのに適しており、疫学的研究のより堅牢な基盤を提供する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。