Skip to main content
QUICK REVIEW

[논문 리뷰] Evaluation of GPT and BERT-based models on identifying protein-protein interactions in biomedical text

Hasin Rehana, Nur Bengisu Çam|PubMed|2023. 03. 30.
Biomedical Text Mining and Ontologies참고 문헌 43인용 수 20
한 줄 요약

본 논문은 GPT 및 BERT 기반 모델을 단백질-단백질 상호작용(PPI) 추출에 대해 세 가지 골드 표준 말뭉치에서 평가하여, BERT 변형이 overall에서 최적의 성능을 보이고 GPT-4는 경쟁력 있는 결과를 보임을 확인한다.

ABSTRACT

Detecting protein-protein interactions (PPIs) is crucial for understanding genetic mechanisms, disease pathogenesis, and drug design. However, with the fast-paced growth of biomedical literature, there is a growing need for automated and accurate extraction of PPIs to facilitate scientific knowledge discovery. Pre-trained language models, such as generative pre-trained transformers (GPT) and bidirectional encoder representations from transformers (BERT), have shown promising results in natural language processing (NLP) tasks. We evaluated the performance of PPI identification of multiple GPT and BERT models using three manually curated gold-standard corpora: Learning Language in Logic (LLL) with 164 PPIs in 77 sentences, Human Protein Reference Database with 163 PPIs in 145 sentences, and Interaction Extraction Performance Assessment with 335 PPIs in 486 sentences. BERT-based models achieved the best overall performance, with BioBERT achieving the highest recall (91.95%) and F1-score (86.84%) and PubMedBERT achieving the highest precision (85.25%). Interestingly, despite not being explicitly trained for biomedical texts, GPT-4 achieved commendable performance, comparable to the top-performing BERT models. It achieved a precision of 88.37%, a recall of 85.14%, and an F1-score of 86.49% on the LLL dataset. These results suggest that GPT models can effectively detect PPIs from text data, offering promising avenues for application in biomedical literature mining. Further research could explore how these models might be fine-tuned for even more specialized tasks within the biomedical domain.

연구 동기 및 목표

  • 생물의학 텍스트에서 PPI를 식별하기 위한 GPT 및 BERT 기반 모델의 효과를 평가한다.
  • 여러 개의 잘 선별된 PPI 코퍼스 간의 성능을 비교한다.
  • PPI 추출 작업에서 최고 정밀도, 재현율, 및 F1를 달성하는 모델을 식별한다.

제안 방법

  • PPI 식별을 평가하기 위해 세 개의 수작업으로 큐레이션된 골드 표준 코퍼라(LLL, HPRD, IPEA)를 사용한다.
  • PPI 작업에서 다수의 GPT 및 BERT 기반 모델을 벤치마크한다.
  • 각 모델과 코퍼스에 대해 정밀도, 재현율, 및 F1-스코어를 보고한다.
  • 최상위 성능 모델(BioBERT, PubMedBERT)과 GPT-4의 상대적 성능을 강조한다.

실험 결과

연구 질문

  • RQ1GPT 기반 모델(GPT-4 포함)이 생물의학 텍스트에서 PPI를 식별하는 정도가 BERT 기반 모델에 비해 얼마나 우수한가?
  • RQ2세 가지 골드 표준 코퍼스에서 어떤 모델 변형이 최상의 정밀도, 재현율, 및 F1를 달성하는가?
  • RQ3생물의학 분야에서 사전 학습되지 않은 GPT 모델들(예: GPT-4)이 PPI 추출에서 전문 생물의학 BERT 모델과 대등한가?
  • RQ4다른 데이터셋에서 BioBERT, PubMedBERT, 및 GPT-4 간 PPI 탐지 성능의 트레이드오프는 무엇인가?

주요 결과

  • BERT 기반 모델이 코퍼스 전반에서 최고 성능을 달성했다.
  • BioBERT가 재현율 91.95% 및 F1-스코어 86.84%로 가장 높았다.
  • PubMedBERT가 정밀도 85.25%로 가장 높았다.
  • GPT-4, 생물의학 텍스트에 특별히 학습되지 않았음에도 불구하고 우수한 성능을 보여 상위 BERT 모델과 견줄 만했다.
  • LLL 데이터셋에서 GPT-4의 정밀도 88.37%, 재현율 85.14%, 및 F1 86.49%를 달성했다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.