Skip to main content
QUICK REVIEW

[论文解读] Evaluation of large language models for discovery of gene set function

Mengzhou Hu, Sahar Alkhairy|arXiv (Cornell University)|Sep 7, 2023
Bioinformatics and Genomic Networks被引用 8
一句话总结

论文对五种大型语言模型在发现基因集功能方面进行基准测试,结果显示GPT-4能够可靠地识别经过整理的功能和来自组学的新的、可验证的功能及其支持证据,而其他模型的信心度有限或存在误导性。

ABSTRACT

Gene set analysis is a mainstay of functional genomics, but it relies on curated databases of gene functions that are incomplete. Here we evaluate five Large Language Models (LLMs) for their ability to discover the common biological functions represented by a gene set, substantiated by supporting rationale, citations and a confidence assessment. Benchmarking against canonical gene sets from the Gene Ontology, GPT-4 confidently recovered the curated name or a more general concept (73% of cases), while benchmarking against random gene sets correctly yielded zero confidence. Gemini-Pro and Mixtral-Instruct showed ability in naming but were falsely confident for random sets, whereas Llama2-70b had poor performance overall. In gene sets derived from 'omics data, GPT-4 identified novel functions not reported by classical functional enrichment (32% of cases), which independent review indicated were largely verifiable and not hallucinations. The ability to rapidly synthesize common gene functions positions LLMs as valuable 'omics assistants.

研究动机与目标

  • 评估大型语言模型在发现由基因集合代表的常见生物学功能方面的能力。
  • 评估LLMs是否提供支持性理由、引用以及信心评估。
  • 将LLM的表现与Gene Ontology的经典基因集合和随机基因集合进行比较。

提出的方法

  • 在基因集功能发现任务中对五种大型语言模型进行基准测试。
  • 衡量恢复经过整理的GO术语或通用概念的能力。
  • 评估每个模型产生的信心、支持性理由和引用。
  • 在Gene Ontology的经典基因集合上进行测试。
  • 在来自组学数据的基因集合上测试以识别新功能。
  • 评估各模型在错误信心和幻觉风险方面的表现。

实验结果

研究问题

  • RQ1LLMs是否能够恢复对应于Gene Ontology术语或概念的经过整理的基因集合功能?
  • RQ2LLMs是否为发现的功能提供可信的支持性理由和引用?
  • RQ3在随机基因集合上,LLMs在信心和准确性方面的表现如何?
  • RQ4LLMs能否从组学派生的基因集合中识别出超越经典富集的新颖、可验证的功能?
  • RQ5GPT-4、Gemini-Pro、Mixtral-Instruct、Llama2-70b等模型在本任务中的相对优劣是什么?

主要发现

  • GPT-4在基于GO的案例中,能够记住经过整理的名称或更一般的概念,比例为73%。
  • 在与随机基因集合进行基准测试时未给出信心。
  • Gemini-Pro和Mixtral-Instruct能够给出功能名称,但对随机集合存在错误自信。
  • Llama2-70b总体表现较差。
  • GPT-4在组学派生的基因集合中辨识出新颖功能,比例为32%,大多可验证且非幻觉,经过独立评审确认。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。