Skip to main content
QUICK REVIEW

[Paper Review] Evaluation of large language models for discovery of gene set function

Mengzhou Hu, Sahar Alkhairy|arXiv (Cornell University)|Sep 7, 2023
Bioinformatics and Genomic Networks8 citations
TL;DR

The paper benchmarks five LLMs for discovering gene set functions, finding GPT-4 reliably identifies curated functions and novel omics-derived functions with verifiable support, while other models show limited or misleading confidence.

ABSTRACT

Gene set analysis is a mainstay of functional genomics, but it relies on curated databases of gene functions that are incomplete. Here we evaluate five Large Language Models (LLMs) for their ability to discover the common biological functions represented by a gene set, substantiated by supporting rationale, citations and a confidence assessment. Benchmarking against canonical gene sets from the Gene Ontology, GPT-4 confidently recovered the curated name or a more general concept (73% of cases), while benchmarking against random gene sets correctly yielded zero confidence. Gemini-Pro and Mixtral-Instruct showed ability in naming but were falsely confident for random sets, whereas Llama2-70b had poor performance overall. In gene sets derived from 'omics data, GPT-4 identified novel functions not reported by classical functional enrichment (32% of cases), which independent review indicated were largely verifiable and not hallucinations. The ability to rapidly synthesize common gene functions positions LLMs as valuable 'omics assistants.

Motivation & Objective

  • Assess the ability of large language models to discover common biological functions represented by a gene set.
  • Evaluate whether LLMs provide supporting rationale, citations, and a confidence assessment.
  • Compare LLM performance against canonical gene sets from Gene Ontology and random gene sets.

Proposed method

  • Benchmarking five large language models on gene set function discovery tasks.
  • Measure ability to recover curated GO terms or general concepts.
  • Assess confidence, supporting rationale, and citations produced by each model.
  • Test on canonical gene sets from Gene Ontology.
  • Test on gene sets derived from omics data to identify novel functions.
  • Evaluate false confidence and hallucination risk across models.

Experimental results

Research questions

  • RQ1Can LLMs recover curated gene set functions corresponding to Gene Ontology terms or concepts?
  • RQ2Do LLMs provide credible supporting rationale and citations for discovered functions?
  • RQ3How do LLMs perform on random gene sets in terms of confidence and accuracy?
  • RQ4Can LLMs identify novel, verifiable functions from omics-derived gene sets beyond classical enrichment?
  • RQ5What are the comparative strengths and weaknesses of GPT-4, Gemini-Pro, Mixtral-Instruct, Llama2-70b, and other models in this task?

Key findings

  • GPT-4 recovered the curated name or a more general concept in 73% of GO-based cases.
  • Zero confidence was yielded when benchmarking against random gene sets.
  • Gemini-Pro and Mixtral-Instruct could name functions but were falsely confident for random sets.
  • Llama2-70b showed poor overall performance.
  • GPT-4 identified novel functions in omics-derived gene sets in 32% of cases, largely verifiable and not hallucinations, per independent review.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.