Skip to main content
QUICK REVIEW

[論文レビュー] Unsupervised Extraction of Phenotypes from Cancer Clinical Notes for Association Studies

Stefan G. Stark, Stephanie L. Hyland|arXiv (Cornell University)|Apr 29, 2019
Biomedical Text Mining and Ontologies参考文献 31被引用数 4
ひとこと要約

本論文は、医療用語と文のクラスタリングを用いて、構造化されていないがん臨床ノートから表型特徴を非教師ありで抽出する手法を提案する。この手法により、体細胞突然変異プロファイルとの関連解析が可能になる。65,000件の文書と320万件の文を対象に適用した結果、341件の有意な関連が同定され、そのうち32件は新規で生物学的に妥当な仮説であり、臨床的特徴と遺伝子突然変異を結びつけるものである。

ABSTRACT

The recent adoption of Electronic Health Records (EHRs) by health care providers has introduced an important source of data that provides detailed and highly specific insights into patient phenotypes over large cohorts. These datasets, in combination with machine learning and statistical approaches, generate new opportunities for research and clinical care. However, many methods require the patient representations to be in structured formats, while the information in the EHR is often locked in unstructured texts designed for human readability. In this work, we develop the methodology to automatically extract clinical features from clinical narratives from large EHR corpora without the need for prior knowledge. We consider medical terms and sentences appearing in clinical narratives as atomic information units. We propose an efficient clustering strategy suitable for the analysis of large text corpora and to utilize the clusters to represent information about the patient compactly. To demonstrate the utility of our approach, we perform an association study of clinical features with somatic mutation profiles from 4,007 cancer patients and their tumors. We apply the proposed algorithm to a dataset consisting of about 65 thousand documents with a total of about 3.2 million sentences. We identify 341 significant statistical associations between the presence of somatic mutations and clinical features. We annotated these associations according to their novelty, and report several known associations. We also propose 32 testable hypotheses where the underlying biological mechanism does not appear to be known but plausible. These results illustrate that the automated discovery of clinical features is possible and the joint analysis of clinical and genetic datasets can generate appealing new hypotheses.

研究の動機と目的

  • 大規模な電子健康記録(EHR)データセットにおける構造化されていない臨床ナラティブから、実用的な表型情報を取り出すという課題に対処すること。
  • 事前の知識や手動による特徴ラベル付けを必要としない、スケーラブルな非教師あり手法を開発すること。
  • 抽出された臨床的特徴とがん患者の体細胞突然変異プロファイルとの関連解析を可能にすること。
  • 臨床的表型と腫瘍ゲノムの間に新しい、生物学的に妥当な仮説を発見すること。
  • EHRテキストを用いた自動的・大規模なフェノームワイド関連解析の実現可能性を示すこと。

提案手法

  • 本手法は、臨床ノート内の個々の医療用語と文を分析のための原子的情報単位として扱う。
  • 65,000件の臨床文書からなる大規模コーパスにおいて、類似した用語と文を効率的にクラスタリングする戦略を適用する。
  • クラスタは、患者レベルの臨床的特徴を効果的に表現するものとして、構造化された表型プロファイルを形成する。
  • 自然言語処理と埋め込み技術を活用し、教師なしで意味的に類似した臨床表現をグループ化する。
  • 4,007名のがん患者において、各クラスタ(臨床的特徴として)の存在と体細胞突然変異状態との間で関連性テストを実施する。
  • 統計的有意性を評価し、新規または生物学的に妥当な関連は、今後の調査の対象としてマークする。

実験結果

リサーチクエスチョン

  • RQ1臨床ノートの用語と文の非教師ありクラスタリングは、構造化されていないEHRノートから意味のある表型特徴を効果的に抽出できるか?
  • RQ2抽出された臨床的特徴と既知の体細胞突然変異関連との重複度はどの程度か?
  • RQ3本手法は、臨床的表型と腫瘍ゲノムの間に新しい、生物学的に妥当な仮説を生成できるか?
  • RQ4数百万の文を含む大規模なEHRコーパスに適用した場合、このアプローチのスケーラビリティはどの程度か?
  • RQ5非教師あり特徴抽出は、がんゲノム研究におけるフェノームワイド関連解析をどの程度支援できるか?

主な発見

  • 本手法は、4,007名のがん患者において、臨床的特徴と体細胞突然変異との間に341件の統計的に有意な関連を効果的に抽出した。
  • 341件の関連のうち32件は新規であり、生物学的に妥当な仮説として特定され、臨床的観察と腫瘍ゲノムの間の潜在的な新たな関連を示唆している。
  • 本手法は、約65,000件の臨床文書と320万件の文を処理することで、スケーラビリティを実証した。
  • 既知の関連が再現されたため、本手法が確立された臨床-ゲノム関係を検出できる能力を有していることが妥当性を保証した。
  • クラスタリング戦略により、構造化されていないテキストから患者の表型をコンactかつ解釈可能な形で表現できた。
  • 結果として、非教師ありNLP技術ががん研究における大規模かつ仮説生成型の関連解析を効果的に支援できることを示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。