Skip to main content
QUICK REVIEW

[논문 리뷰] Unsupervised Extraction of Phenotypes from Cancer Clinical Notes for Association Studies

Stefan G. Stark, Stephanie L. Hyland|arXiv (Cornell University)|2019. 04. 29.
Biomedical Text Mining and Ontologies참고 문헌 31인용 수 4
한 줄 요약

이 논문은 구조화되지 않은 암 임상 노트에서 의료 용어와 문장의 클러스터링을 통해 비지도 학습 방법을 제안하며, 체계적 돌연변이 프로필과의 연관성 연구를 가능하게 한다. 65,000건의 문서와 320만 개의 문장에 적용된 이 방법은 341개의 유의미한 연관성을 식별하였으며, 이 중 32개는 새로운 생물학적으로 타당한 가설로, 임상적 특징과 유전자 돌연변이를 연결짓는 데 기여한다.

ABSTRACT

The recent adoption of Electronic Health Records (EHRs) by health care providers has introduced an important source of data that provides detailed and highly specific insights into patient phenotypes over large cohorts. These datasets, in combination with machine learning and statistical approaches, generate new opportunities for research and clinical care. However, many methods require the patient representations to be in structured formats, while the information in the EHR is often locked in unstructured texts designed for human readability. In this work, we develop the methodology to automatically extract clinical features from clinical narratives from large EHR corpora without the need for prior knowledge. We consider medical terms and sentences appearing in clinical narratives as atomic information units. We propose an efficient clustering strategy suitable for the analysis of large text corpora and to utilize the clusters to represent information about the patient compactly. To demonstrate the utility of our approach, we perform an association study of clinical features with somatic mutation profiles from 4,007 cancer patients and their tumors. We apply the proposed algorithm to a dataset consisting of about 65 thousand documents with a total of about 3.2 million sentences. We identify 341 significant statistical associations between the presence of somatic mutations and clinical features. We annotated these associations according to their novelty, and report several known associations. We also propose 32 testable hypotheses where the underlying biological mechanism does not appear to be known but plausible. These results illustrate that the automated discovery of clinical features is possible and the joint analysis of clinical and genetic datasets can generate appealing new hypotheses.

연구 동기 및 목표

  • 대규모 전자 건강기록(EHR) 데이터셋에서 비구조화된 임상 서술문에서 실용적인 형상 정보를 추출하는 데 도전하는 것.
  • 사전 지식이나 임상 특징의 수동 레이블링이 필요 없는 확장 가능한 비지도 학습 방법을 개발하는 것.
  • 추출된 임상 특징과 암 환자의 체계적 돌연변이 프로필 간의 연관성 연구를 가능하게 하는 것.
  • 임상 형상와 종양 게놈학을 연결짓는 새로운 생물학적으로 타당한 가설을 발견하는 것.
  • EHR 텍스트를 활용한 자동화되고 대규모의 형상 전체 연관성 연구의 실현 가능성을 입증하는 것.

제안 방법

  • 이 방법은 임상 노트 내 개별 의료 용어와 문장을 분석을 위한 원자적 정보 단위로 간주한다.
  • 65,000건의 임상 문서로 구성된 대규모 코퍼스에서 유사한 용어와 문장을 그룹화하기 위해 효율적인 클러스터링 전략을 적용한다.
  • 클러스터는 환자 수준의 임상 특징을 압축적으로 표현하며, 구조화된 형상 프로파일을 형성한다.
  • 자연어 처리와 임bedding 기법을 활용하여 감독 없이 의미적으로 유사한 임상 표현을 그룹화한다.
  • 각 클러스터(임상 특징으로서)의 존재 여부와 4,007명의 암 환자에서의 체계적 돌연변이 상태 간 연관성 테스트를 수행한다.
  • 통계적 유의성은 평가되며, 새로운 또는 생물학적으로 타당한 연관성은 향후 조사 대상으로 표시된다.

실험 결과

연구 질문

  • RQ1비지도 클러스터링을 통해 임상 텍스트 용어와 문장을 효과적으로 비구조화된 EHR 노트에서 의미 있는 형상 특징으로 추출할 수 있는가?
  • RQ2추출된 임상 특징과 기존의 체계적 돌연변이 연관성 간의 중복 범위는 어느 정도인가?
  • RQ3이 방법은 임상 형상와 종양 게놈학을 연결짓는 새로운 생물학적으로 타당한 가설을 생성할 수 있는가?
  • RQ4수백만 개의 문장이 포함된 대규모 EHR 코퍼스에 적용했을 때 이 방법의 확장성은 어느 정도인가?
  • RQ5비지도 특징 추출 기법이 암 게놈학에서 형상 전체 연관성 연구를 어느 정도 지원할 수 있는가?

주요 결과

  • 이 방법은 4,007명의 암 환자에서 임상 특징과 체계적 돌연변이 간에 341개의 통계적으로 유의미한 연관성을 성공적으로 추출하였다.
  • 341개의 연관성 중 32개는 새로운 것으로 생물학적으로 타당한 것으로 확인되어 임상 관찰과 종양 게놈학 간 잠재적 연결 고리를 시사한다.
  • 약 65,000건의 임상 문서와 320만 개의 문장을 처리함으로써 이 방법의 확장성을 입증하였다.
  • 기존에 알려진 연관성들이 복원되어, 이 방법이 기존의 임상-유전체 관계를 탐지할 수 있는 능력을 입증하였다.
  • 클러스터링 전략은 비구조화된 텍스트에서 환자 형상의 압축되고 해석 가능한 표현을 가능하게 하였다.
  • 결과적으로 비지도 NLP 기법이 대규모, 가설 생성을 위한 연관성 연구에 효과적으로 기여할 수 있음을 보여주었다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.