[논문 리뷰] Unsupervised Machine Learning for the Discovery of Latent Disease Clusters and Patient Subgroups Using Electronic Health Records
이 논문은 전자 건강 기록(EHR)에서 잠재적 질병 군집과 환자 하위군을 비지도 기계학습을 사용하여 규명하기 위해 새로운 푸아송 딜리클 모델(Poisson Dirichlet Model, PDM)을 제안한다. 표준 LDA와 달리 PDM는 진단을 포isson 분포로 모델링하고, 관찰된 진단과 기대 진단을 비교하여 연령과 성별을 보정함으로써 생물학적으로 의미 있는 공존 질환 패턴을 드러내며, 인구통계적 편향을 감소시킨다.
Machine learning has become ubiquitous and a key technology on mining electronic health records (EHRs) for facilitating clinical research and practice. Unsupervised machine learning, as opposed to supervised learning, has shown promise in identifying novel patterns and relations from EHRs without using human created labels. In this paper, we investigate the application of unsupervised machine learning models in discovering latent disease clusters and patient subgroups based on EHRs. We utilized Latent Dirichlet Allocation (LDA), a generative probabilistic model, and proposed a novel model named Poisson Dirichlet Model (PDM), which extends the LDA approach using a Poisson distribution to model patients' disease diagnoses and to alleviate age and sex factors by considering both observed and expected observations. In the empirical experiments, we evaluated LDA and PDM on three patient cohorts with EHR data retrieved from the Rochester Epidemiology Project (REP), for the discovery of latent disease clusters and patient subgroups. We compared the effectiveness of LDA and PDM in identifying latent disease clusters through the visualization of disease representations learned by two approaches. We also tested the performance of LDA and PDM in differentiating patient subgroups through survival analysis, as well as statistical analysis. The experimental results show that the proposed PDM could effectively identify distinguished disease clusters by alleviating the impact of age and sex, and that LDA could stratify patients into more differentiable subgroups than PDM in terms of p-values. However, the subgroups discovered by PDM might imply the underlying patterns of diseases of greater interest in epidemiology research due to the alleviation of age and sex. Both unsupervised machine learning approaches could be leveraged to discover patient subgroups using EHRs but with different foci.
연구 동기 및 목표
- 라벨이 없는 결과에 의존하지 않고 EHR 데이터에서 잠재적 질병 군집과 환자 하위군을 식별하는 것.
- 노령 인구에서 흔히 관찰되는 공존 질환 패턴에 영향을 미치는 연령과 성별의 혼동 요인을 해결하는 것.
- 기대 진단 수를 통합함으로써 LDA를 확장한 새로운 비지도 모델인 푸아송 딜리클 모델(Poisson Dirichlet Model, PDM)을 개발하고 평가하는 것.
- 실제 EHR 데이터를 사용하여 PDM과 LDA의 성능을 비교하여 생물학적으로 의미 있는 질병 군집과 환자 하위군을 발견하는 능력을 평가하는 것.
- 생존 분석과 공존 질환 프로파일링을 통해 발견된 하위군의 임상적 및 역학적 관련성을 평가하는 것.
제안 방법
- 질병 군집을 진단에 대한 확률 분포로 표현하는 비지도 주제 모델로 기반을 삼기 위해 수정된 은닉 딜리클 할당(Latent Dirichlet Allocation, LDA)을 사용.
- 계수 기반의 진단 데이터를 다룰 수 있도록 환자의 진단을 포isson 분포로 모델링하는 새로운 푸아송 딜리클 모델(Poisson Dirichlet Model, PDM)을 제안.
- 연령-성별 그룹별로 관찰된 진단 수와 기대 진단 수를 모두 통합하여 질병 유병률의 인구통계적 편향을 보정.
- 잠재적 질병 군집의 확률적 추론을 가능하게 하기 위해 질병-주제 분포에 딜리클 우선사전을 적용.
- LDA와 PDM의 모든 파라미터 추정에 마르코프 체인 몬테카를로(Markov Chain Monte Carlo, MCMC) 방법을 적용.
- LDA와 PDM 간의 군집 분리 정도를 비교하기 위해 2차원 잠재 주제 공간에서 질병 표현을 시각화.
실험 결과
연구 질문
- RQ1제안된 PDM은 연령과 성별의 혼동 효과를 최소화하면서 EHR 데이터에서 명확한 잠재 질병 군집을 효과적으로 식별할 수 있는가?
- RQ2PDM이 식별한 환자 하위군은 LDA에 의해 식별된 하위군과 임상적 및 인구통계적 특성에서 어떻게 다를까?
- RQ3PDM이 발견한 환자 하위군은 연령과 성별을 초월해 생물학적 또는 역학적으로 의미 있는 패턴을 반영하는가?
- RQ4LDA와 PDM은 생존 분석에서 환자를 얼마나 잘 분류하는가? 이는 예후의 차별화를 의미하는가?
- RQ5각 모델이 식별한 하위군 간에 엘렉시아우어 공존 질환 지수(ECI) 점수는 어느 정도 다를까?
주요 결과
- PDM는 연령과 성별에 기반한 기대 질병 유병률을 고려함으로써 LDA보다 더 명확하고 생물학적으로 해석 가능한 질병 군집을 성공적으로 식별했다.
- 생존 분석에서 LDA는 PDM보다 환자 하위군을 더 효과적으로 분류하여 유의미한 p값(p < 0.001)을 기록함으로써 통계적 분리 능력이 뛰어났다.
- 생존 분석에서 통계적 유의성이 낮았음에도 불구하고, PDM이 도출한 하위군은 연령과 성별의 혼동 영향을 줄여 더 높은 역학적 관련성을 보였다.
- PDM이 식별한 하위군 간에 엘렉시아우어 공존 질환 지수(ECI) 점수가 유의미하게 다름을 보여, 공존 질환 부담의 임상적 다양성이 뚜렷하게 드러났다.
- 질병 표현의 시각화 결과, PDM은 LDA보다 잠재 주제 공간에서 더 명확하고 더 극단적인 군집을 생성했다.
- 제안된 PDM 모델은 연령과 성별이 주요 혼동 요인인 노령 인구에서 잠재적 질병 패턴을 식별하는 데 더 적합하며, 역학 연구에 더 견고한 기반을 제공한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.