[논문 리뷰] Comparative Performance Evaluation of Large Language Models for Extracting Molecular Interactions and Pathway Knowledge
이 논문은 bioNLP 과제에서 여러 대형 언어 모델(LLMs)을 체계적으로 비교합니다: 단백질–단백질 상호작용, 저선량 방사선에 의해 영향을 받는 경로, 및 유전자 조절 관계를 다루며, 설정 및 데이터 소스 전반에서 어떤 모델이 최고 성능을 내는지 식별합니다.
Background: Identification of the interactions and regulatory relations between biomolecules play pivotal roles in understanding complex biological systems and the mechanisms underlying diverse biological functions. However, the collection of such molecular interactions has heavily relied on expert curation in the past, making it labor-intensive and time-consuming. To mitigate these challenges, we propose leveraging the capabilities of large language models (LLMs) to automate genome-scale extraction of this crucial knowledge. Results: In this study, we investigate the efficacy of various LLMs in addressing biological tasks, such as the recognition of protein interactions, identification of genes linked to pathways affected by low-dose radiation, and the delineation of gene regulatory relationships. Overall, the larger models exhibited superior performance, indicating their potential for specific tasks that involve the extraction of complex interactions among genes and proteins. Although these models possessed detailed information for distinct gene and protein groups, they faced challenges in identifying groups with diverse functions and in recognizing highly correlated gene regulatory relationships. Conclusions: By conducting a comprehensive assessment of the state-of-the-art models using well-established molecular interaction and pathway databases, our study reveals that LLMs can identify genes/proteins associated with pathways of interest and predict their interactions to a certain extent. Furthermore, these models can provide important insights, marking a noteworthy stride toward advancing our understanding of biological systems through AI-assisted knowledge discovery.
연구 동기 및 목표
- 생물의학 문헌에서 분자 상호작용과 경로 지식 추출에서 다양한 대형 언어 모델(LLMs)의 효과를 평가한다.
- PPI 인식, KEGG 데이터를 이용한 LDR 영향 경로 유전자 회수, 및 다수의 LLM 간 유전자 조절 관계 태스크를 비교한다.
- 특정 생물학적 지식 추출 태스크에 대해 어떤 모델이 뛰어난지 식별하고 한계와 기회를 논의한다.
제안 방법
- Galactica, Alpaca, RST, Falcon, MPT, LLaMA2 및 도메인 전문 BioGPT/BioMedLM을 포함한 여러 LLM을 세 bioNLP 태스크에서 평가한다.
- 데이터 소스로 STRING, KEGG, 및 INDRA를 사용하여 PPI, 경로 유전자, 및 유전자 조절 관계에 대한 평가 세트를 구성한다.
- 작업당 최적의 프롬프트 전략을 식별하기 위해 맥락 예시 수(0–5 샷)와 프롬프트를 다양화한다.
- 작업별 배치 크기로 4× NVIDIA A100 80GB GPU에서 실험을 수행한다.
- 성능을 정량화하기 위해 micro F1, macro F1, 및 전체 매치 수를 보고한다.

실험 결과
연구 질문
- RQ1STRING 유래 인간 단백질 네트워크에서 어떤 LLM이 단백질–단백질 상호작용을 가장 잘 인식하는가?
- RQ2KEGG 데이터를 사용하여 저선량 방사선(LDR) 노출에 영향을 받은 인간 경로의 유전자를 가장 정확하게 식별하는 모델은 무엇인가?
- RQ3INDRA DB 텍스트 진술을 사용하여 LLM이 유전자 조절 관계를 얼마나 잘 분류하는가?
- RQ4모델 크기나 도메인 전문화가 태스크 간 성능과 상관관계가 있는가?
- RQ5프롬프트 전략(샷)이 각 태스크의 성능에 어떤 영향을 미치는가?
주요 결과
- LLaMA2-Chat (70B)가 PPI Task1에서 가장 높은 Micro F1 및 Macro F1을 달성했으며 1K 중 159건의 전체 일치가 나옴.
- LLaMA2-Chat (7B)는 PPI Task1에서 MPT-Chat (30B) 및 Galactica (30B)와 근접한 성능을 보임.
- MPT-Chat (7B) 및 MPT-Chat (30B)가 PPI Task2(이진 예/아니오)에서 가장 강하게 perform하며 Micro F1이 각각 최대 0.9840 및 0.9350에 도달(5샷).
- LDR 경로 태스크에서 Galactica (30B)와 MPT-Chat (30B)가 가장 정확하게 유전자를 예측하고, BioMedLM 및 BioGPT-Large가 도메인 특화 데이터에서 주목할 만한 이점을 보임.
- INDRA 태스크에서 더 큰 모델(예: Galactica 30B, LLaMA2-Chat 70B, MPT-Chat 30B)이 더 작은 BioGPT/BioMedLM 모델보다 우수하여 유전자 조절 관계에 대한 독해에 크기와 다양한 학습 데이터가 도움을 줌.
- 도메인 특화된 소형 모델이 특정 전문 태스크에서 더 큰 일반 모델을 능가할 수 있어 태스크-도메인 정렬의 중요성을 시사함.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.