[논문 리뷰] Using Machine Learning and Natural Language Processing to Review and Classify the Medical Literature on Cancer Susceptibility Genes
이 연구는 생물의학 문헌 초록을 암 양성유전성 유전자 침습성 또는 유병률과 관련이 있는지 자동으로 분류하기 위해 지원벡터기계(SVM)와 합성곱신경망(CNN)이라는 두 가지 기계학습 모델을 개발하고 평가한다. SVM는 침입성 분류에서 89.53%의 정확도를 기록했으며, 유병률 분류에서는 89.14%의 정확도를 기록하여 임상 유전학 분야에서 확장 가능한 문헌 검토에 높은 성능을 보였다.
PURPOSE: The medical literature relevant to germline genetics is growing exponentially. Clinicians need tools monitoring and prioritizing the literature to understand the clinical implications of the pathogenic genetic variants. We developed and evaluated two machine learning models to classify abstracts as relevant to the penetrance (risk of cancer for germline mutation carriers) or prevalence of germline genetic mutations. METHODS: We conducted literature searches in PubMed and retrieved paper titles and abstracts to create an annotated dataset for training and evaluating the two machine learning classification models. Our first model is a support vector machine (SVM) which learns a linear decision rule based on the bag-of-ngrams representation of each title and abstract. Our second model is a convolutional neural network (CNN) which learns a complex nonlinear decision rule based on the raw title and abstract. We evaluated the performance of the two models on the classification of papers as relevant to penetrance or prevalence. RESULTS: For penetrance classification, we annotated 3740 paper titles and abstracts and used 60% for training the model, 20% for tuning the model, and 20% for evaluating the model. The SVM model achieves 89.53% accuracy (percentage of papers that were correctly classified) while the CNN model achieves 88.95 % accuracy. For prevalence classification, we annotated 3753 paper titles and abstracts. The SVM model achieves 89.14% accuracy while the CNN model achieves 89.13 % accuracy. CONCLUSION: Our models achieve high accuracy in classifying abstracts as relevant to penetrance or prevalence. By facilitating literature review, this tool could help clinicians and researchers keep abreast of the burgeoning knowledge of gene-cancer associations and keep the knowledge bases for clinical decision support tools up to date.
연구 동기 및 목표
- 급격히 증가하는 유전성 유전학 문헌 문제를 해결하기 위해 관련 초록의 분류를 자동화한다.
- 임상의가 암 양성유전성 유전자 침입성과 유병률에 관한 핵심 연구를 식별하는 데 도움이 되는 기계학습 도구를 개발한다.
- 학습 및 평가를 위한 분류된 데이터셋을 3,740건(침입성)과 3,753건(유병률)으로 구축한다.
- 동일한 분류 작업에 대해 전통적인 SVM와 딥러닝 기반 CNN 모델의 성능을 비교한다.
- 유전성 변이 관련성에 대한 지속적이고 스케일 가능한 지식 기반의 유지보수를 가능하게 한다.
제안 방법
- PubMed 문헌 검색을 통해 유전성 암 양성유전성 유전자와 관련된 제목과 초록을 수집한다.
- 전문가 검토를 통해 3,740건의 초록을 침입성 관련성, 3,753건을 유병률 관련성으로 분류한다.
- 텍스트 특징의 나그네-그램 표현을 사용하여 지원벡터기계(SVM) 모델을 훈련시킨다.
- 원시 텍스트 입력을 사용하여 비선형 결정 경계를 학습하는 합성곱신경망(CNN) 모델을 훈련시킨다.
- 모든 모델에 대해 데이터를 60% 훈련, 20% 하이퍼파ram터 조정, 20% 평가로 분할한다.
- 보류된 테스트 세트에서 정확도를 주요 평가 지표로 사용하여 모델 성능을 평가한다.
실험 결과
연구 질문
- RQ1기계학습 모델은 암 양성유전성 유전자 침입성과 관련된 생물의학 초록을 정확하게 분류할 수 있는가?
- RQ2기계학습 모델은 암 양성유전성 유전자 유병률과 관련된 초록을 정확하게 분류할 수 있는가?
- RQ3기존의 SVM와 딥러닝 기반 CNN 모델은 유전성 변이 문헌 분류 작업에서 어떻게 성능을 비교하는가?
- RQ4자동 분류 기술은 임상 유전학 분야에서 수동 문헌 검토의 부담을 어느 정도 줄일 수 있는가?
- RQ5이러한 모델은 임상 의사결정 지원 도구를 위한 최신 지식 기반 유지에 기여할 수 있는가?
주요 결과
- SVM 모델은 암 양성유전성 유전자 침입성과 관련된 초록 분류에서 89.53%의 정확도를 기록했다.
- CNN 모델은 침입성 분류에서 88.95%의 정확도를 기록했으며, 복잡성에도 불구하고 뛰어난 성능을 보였다.
- 유병률 분류에서는 SVM 모델이 89.14%의 정확도를 기록했고, CNN 모델은 89.13%의 정확도를 달성했다.
- 두 모델 모두 높고 유사한 정확도를 보였으며, 이는 더 단순한 모델인 SVM도 이 작업에 효과적일 수 있음을 시사한다.
- 결과는 자동 분류 도구가 유전성 암 위험에 대한 지속적인 모니터링과 업데이트에 필요한 수동 작업을 크게 줄일 수 있음을 시사한다.
- 이 연구는 자연어 처리와 기계학습을 활용하여 임상 유전학 적용 분야에서 문헌 정제를 확장하는 데의 가능성을 입증한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.