Skip to main content
QUICK REVIEW

[Paper Review] Benchmarking large language models for biomedical natural language processing applications and recommendations

Qingyu Chen, Yan Hu|arXiv (Cornell University)|May 10, 2023
Topic Modeling41 citations
TL;DR

This study systematically evaluates four large language models across 12 BioNLP benchmarks, comparing zero-shot, few-shot, and fine-tuning against traditional BERT/BART fine-tuning, with recommendations.

ABSTRACT

The rapid growth of biomedical literature poses challenges for manual knowledge curation and synthesis. Biomedical Natural Language Processing (BioNLP) automates the process. While Large Language Models (LLMs) have shown promise in general domains, their effectiveness in BioNLP tasks remains unclear due to limited benchmarks and practical guidelines. We perform a systematic evaluation of four LLMs, GPT and LLaMA representatives on 12 BioNLP benchmarks across six applications. We compare their zero-shot, few-shot, and fine-tuning performance with traditional fine-tuning of BERT or BART models. We examine inconsistencies, missing information, hallucinations, and perform cost analysis. Here we show that traditional fine-tuning outperforms zero or few shot LLMs in most tasks. However, closed-source LLMs like GPT-4 excel in reasoning-related tasks such as medical question answering. Open source LLMs still require fine-tuning to close performance gaps. We find issues like missing information and hallucinations in LLM outputs. These results offer practical insights for applying LLMs in BioNLP.

Motivation & Objective

  • Address how LLMs perform on BioNLP tasks relative to traditional models.
  • Assess zero-shot, few-shot, and fine-tuning capabilities of LLMs in biomedical contexts.
  • Identify inconsistencies, missing information, and hallucinations in LLM outputs.
  • Evaluate cost implications of using LLMs in BioNLP applications.
  • Provide practical recommendations for applying LLMs in BioNLP.

Proposed method

  • Systematically evaluate four LLMs (GPT and LLaMA representatives) on 12 BioNLP benchmarks across six applications.
  • Compare zero-shot, few-shot, and fine-tuning performance of LLMs against traditional fine-tuning of BERT or BART models.
  • Analyze output quality for inconsistencies, missing information, and hallucinations.
  • Conduct cost analysis of LLM usage in BioNLP tasks.

Experimental results

Research questions

  • RQ1How do LLMs perform on BioNLP benchmarks in zero-shot, few-shot, and fine-tuned settings compared to traditional BERT/BART fine-tuning?
  • RQ2Are closed-source LLMs (e.g., GPT-4) better for reasoning-related BioNLP tasks such as medical question answering?
  • RQ3To what extent do open-source LLMs require fine-tuning to close performance gaps with traditional models?
  • RQ4What are common issues (missing information, hallucinations) in LLM outputs within BioNLP tasks?
  • RQ5What practical guidelines can be derived for applying LLMs in BioNLP from this benchmarking study?

Key findings

  • Traditional fine-tuning generally outperforms zero- or few-shot LLMs in most BioNLP tasks.
  • Closed-source LLMs like GPT-4 excel in reasoning-related tasks such as medical question answering.
  • Open-source LLMs still require fine-tuning to close performance gaps with traditional models.
  • LLM outputs exhibit missing information and hallucinations that affect reliability in BioNLP.
  • The study provides practical insights and recommendations for applying LLMs in BioNLP.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.