Skip to main content
QUICK REVIEW

[论文解读] Unsupervised Machine Learning for the Discovery of Latent Disease Clusters and Patient Subgroups Using Electronic Health Records

Yanshan Wang, Yiqing Zhao|arXiv (Cornell University)|May 17, 2019
Machine Learning in Healthcare参考文献 32被引用 6
一句话总结

本文提出了一种新颖的泊松狄利克雷模型(PDM),利用无监督机器学习从电子健康记录(EHRs)中揭示潜在的疾病聚集和患者亚群。与标准LDA不同,PDM采用泊松分布对疾病诊断进行建模,并通过比较实际诊断与预期诊断来调整年龄和性别因素,从而揭示具有生物学意义的共病模式,同时减少人口统计学偏差。

ABSTRACT

Machine learning has become ubiquitous and a key technology on mining electronic health records (EHRs) for facilitating clinical research and practice. Unsupervised machine learning, as opposed to supervised learning, has shown promise in identifying novel patterns and relations from EHRs without using human created labels. In this paper, we investigate the application of unsupervised machine learning models in discovering latent disease clusters and patient subgroups based on EHRs. We utilized Latent Dirichlet Allocation (LDA), a generative probabilistic model, and proposed a novel model named Poisson Dirichlet Model (PDM), which extends the LDA approach using a Poisson distribution to model patients' disease diagnoses and to alleviate age and sex factors by considering both observed and expected observations. In the empirical experiments, we evaluated LDA and PDM on three patient cohorts with EHR data retrieved from the Rochester Epidemiology Project (REP), for the discovery of latent disease clusters and patient subgroups. We compared the effectiveness of LDA and PDM in identifying latent disease clusters through the visualization of disease representations learned by two approaches. We also tested the performance of LDA and PDM in differentiating patient subgroups through survival analysis, as well as statistical analysis. The experimental results show that the proposed PDM could effectively identify distinguished disease clusters by alleviating the impact of age and sex, and that LDA could stratify patients into more differentiable subgroups than PDM in terms of p-values. However, the subgroups discovered by PDM might imply the underlying patterns of diseases of greater interest in epidemiology research due to the alleviation of age and sex. Both unsupervised machine learning approaches could be leveraged to discover patient subgroups using EHRs but with different foci.

研究动机与目标

  • 在不依赖标注结果的情况下,从EHR数据中识别潜在的疾病聚集和患者亚群。
  • 解决在老龄化人群中常见的共病模式中年龄和性别造成的混杂影响。
  • 开发并评估一种新颖的无监督模型——泊松狄利克雷模型(PDM),该模型通过引入预期疾病计数,扩展了LDA。
  • 使用真实世界EHR数据,比较PDM与LDA在发现具有生物学意义的疾病聚集和患者亚群方面的性能。
  • 通过生存分析和共病特征分析,评估所发现亚群的临床和流行病学相关性。

提出的方法

  • 将改进的潜在狄利克雷分配(LDA)作为基线无监督主题模型,将疾病聚集表示为诊断的概率分布。
  • 提出一种新颖的泊松狄利克雷模型(PDM),通过泊松分布对患者诊断进行建模,以处理基于计数的诊断数据。
  • 结合每个年龄-性别组的实际和预期疾病计数,以调整疾病流行率中的人口统计学偏差。
  • 在疾病-主题分布上使用狄利克雷先验,以支持潜在疾病聚集的概率推断。
  • 在LDA和PDM中均采用马尔可夫链蒙特卡洛(MCMC)方法进行参数估计。
  • 在二维潜在主题空间中可视化疾病表示,以比较LDA与PDM在聚类分离度上的差异。

实验结果

研究问题

  • RQ1所提出的PDM能否在最小化年龄和性别混杂效应的同时,有效识别EHR数据中不同的潜在疾病聚集?
  • RQ2PDM识别的患者亚群与LDA识别的亚群在临床和人口统计学可区分性方面有何差异?
  • RQ3PDM发现的患者亚群是否反映了超越年龄和性别之外的生物学或流行病学上有意义的模式?
  • RQ4LDA与PDM在生存分析中对患者的分层效果如何,表明其预后区分能力?
  • RQ5各模型识别的亚群之间,Elixhauser共病指数(ECI)得分的差异程度如何?

主要发现

  • PDM通过基于年龄和性别调整预期疾病流行率,成功识别出比LDA更清晰且更具生物学可解释性的疾病聚集。
  • 尽管在生存分析中统计显著性较低,PDM识别的亚群在预后区分上表现更优,其p值低于0.001,表明更强的统计差异。
  • 尽管在生存分析中统计显著性较低,PDM识别的亚群在预后区分上表现更优,其p值低于0.001,表明更强的统计差异。
  • PDM识别的亚群之间Elixhauser共病指数(ECI)得分存在显著差异,提示共病负担存在有意义的临床差异。
  • 疾病表示的可视化显示,PDM在潜在主题空间中产生的聚类比LDA更清晰、更具二元分离性。
  • 所提出的PDM模型更适合在年龄和性别为主要混杂因素的老年队列中识别潜在疾病模式,为流行病学研究提供了更稳健的基础。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。