[論文レビュー] Using Machine Learning and Natural Language Processing to Review and Classify the Medical Literature on Cancer Susceptibility Genes
本研究では、がん感受性遺伝子の発現率や有病率に関連するバイオメディカル文献の要約を自動的に分類するための2つの機械学習モデル—サポートベクターマシン(SVM)および畳み込みニューラルネットワーク(CNN)—を開発および評価した。SVMは発現率分類で89.53%、有病率分類で89.14%の正確性を達成し、臨床遺伝学分野におけるスケーラブルな文献レビューにおいて高い性能を示した。
PURPOSE: The medical literature relevant to germline genetics is growing exponentially. Clinicians need tools monitoring and prioritizing the literature to understand the clinical implications of the pathogenic genetic variants. We developed and evaluated two machine learning models to classify abstracts as relevant to the penetrance (risk of cancer for germline mutation carriers) or prevalence of germline genetic mutations. METHODS: We conducted literature searches in PubMed and retrieved paper titles and abstracts to create an annotated dataset for training and evaluating the two machine learning classification models. Our first model is a support vector machine (SVM) which learns a linear decision rule based on the bag-of-ngrams representation of each title and abstract. Our second model is a convolutional neural network (CNN) which learns a complex nonlinear decision rule based on the raw title and abstract. We evaluated the performance of the two models on the classification of papers as relevant to penetrance or prevalence. RESULTS: For penetrance classification, we annotated 3740 paper titles and abstracts and used 60% for training the model, 20% for tuning the model, and 20% for evaluating the model. The SVM model achieves 89.53% accuracy (percentage of papers that were correctly classified) while the CNN model achieves 88.95 % accuracy. For prevalence classification, we annotated 3753 paper titles and abstracts. The SVM model achieves 89.14% accuracy while the CNN model achieves 89.13 % accuracy. CONCLUSION: Our models achieve high accuracy in classifying abstracts as relevant to penetrance or prevalence. By facilitating literature review, this tool could help clinicians and researchers keep abreast of the burgeoning knowledge of gene-cancer associations and keep the knowledge bases for clinical decision support tools up to date.
研究の動機と目的
- 急速に増加するゲノム遺伝学分野の文献に対応するため、関連する要約の自動分類を実現する。
- 臨床医ががん感受性遺伝子の発現率および有病率に関する重要な研究を特定するのを支援する機械学習ツールを開発する。
- 訓練および評価用の分類モデルのため、3,740件(発現率)および3,753件(有病率)の要約から構成されるラベル付きデータセットを作成する。
- 同じ分類タスクにおいて、従来のSVMとディープラーニングのCNNモデルの性能を比較する。
- 遺伝性がんリスク評価のための臨床意思決定支援システムで使用される知識ベースのスケーラブルかつ最新の保守を可能にする。
提案手法
- PubMedを用いて、ゲノムがん感受性遺伝子に関連するタイトルおよび要約を収集した。
- 専門家によるレビューを用いて、3,740件の要約を発現率関連、3,753件を有病率関連としてラベル付けした。
- テキスト特徴量のbag-of-ngrams表現を用いて、サポートベクターマシン(SVM)モデルを訓練した。
- 生のテキスト入力を用いて、非線形の意思決定境界を学習する畳み込みニューラルネットワーク(CNN)モデルを訓練した。
- 両モデルのため、60%を訓練、20%をハイパーパramータチューニング、20%を評価用にデータを分割した。
- 保持されたテストセット上で、正確性を主な指標としてモデルの性能を評価した。
実験結果
リサーチクエスチョン
- RQ1機械学習モデルは、がん感受性遺伝子の発現率に関連するバイオメディカル要約を正確に分類できるか?
- RQ2機械学習モデルは、がん感受性遺伝子の有病率に関連する要約を正確に分類できるか?
- RQ3従来のSVMとディープラーニングのCNNモデルは、ゲノムバリアント関連文献の分類において、性能でどのように比較できるか?
- RQ4自動分類は、臨床遺伝学分野における手動の文献レビューの負担をどの程度軽減できるか?
- RQ5これらのモデルは、臨床意思決定支援ツールのための最新の知識ベースの保守を支援できるか?
主な発見
- SVMモデルは、がん感受性遺伝子の発現率に関連する要約の分類において89.53%の正確性を達成した。
- CNNモデルは発現率分類で88.95%の正確性を達成し、その複雑さにもかかわらず優れた性能を示した。
- 有病率分類においては、SVMモデルが89.14%の正確性を達成し、CNNモデルは89.13%の正確性を達成した。
- 両モデルともに高いかつ同等の正確性を示しており、SVMのような単純なモデルがこのタスクにおいても有効である可能性を示している。
- 結果から、自動分類ツールが、ゲノムがんリスクに関する知識の監視および更新に必要な手作業の負担を顕著に軽減できることが示唆された。
- 本研究では、自然言語処理および機械学習を活用して、臨床遺伝学の応用分野における文献のキュレートをスケーラブルに実現する可能性が示された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。