Skip to main content
QUICK REVIEW

[论文解读] LitSumm: Large language models for literature summarisation of non-coding RNAs

Andrew R. Green, Carlos Eduardo Ribas|arXiv (Cornell University)|Nov 6, 2023
RNA modifications and cancer被引用 4
一句话总结

本文介绍了 LitSumm 系统,该系统利用微调后的大型语言模型(LLMs)结合结构化提示和自动化验证,生成高质量、事实准确的非编码RNA(ncRNA)文献摘要。该方法在人工评估中得分较高,并已应用于超过4,600个ncRNA,相关摘要已通过RNAcentral公开提供。

ABSTRACT

Curation of literature in life sciences is a growing challenge. The continued increase in the rate of publication, coupled with the relatively fixed number of curators worldwide presents a major challenge to developers of biomedical knowledgebases. Very few knowledgebases have resources to scale to the whole relevant literature and all have to prioritise their efforts. In this work, we take a first step to alleviating the lack of curator time in RNA science by generating summaries of literature for non-coding RNAs using large language models (LLMs). We demonstrate that high-quality, factually accurate summaries with accurate references can be automatically generated from the literature using a commercial LLM and a chain of prompts and checks. Manual assessment was carried out for a subset of summaries, with the majority being rated extremely high quality. We apply our tool to a selection of over 4,600 ncRNAs and make the generated summaries available via the RNAcentral resource. We conclude that automated literature summarization is feasible with the current generation of LLMs, provided careful prompting and automated checking are applied.

研究动机与目标

  • 为应对非编码RNA(ncRNA)研究中文献快速增长带来的数据整理挑战。
  • 通过使用大型语言模型自动化文献摘要,减少对人工整理的依赖。
  • 开发一种可扩展、可靠的摘要生成方法,确保事实准确并正确引用文献。
  • 通过人工评估衡量大型语言模型生成摘要的质量。
  • 实现系统的大规模部署,为超过4,600个ncRNA生成摘要,并将其整合至RNAcentral知识库。

提出的方法

  • 采用商业大型语言模型(LLM),结合思维链提示策略,生成结构化摘要。
  • 设计一系列提示,引导LLM提取关键事实,包括功能角色、表达模式及分子互作关系。
  • 实施自动化检查,验证生成摘要中的事实一致性和引用准确性。
  • 通过引用定位技术,确保摘要中的引用与实际文献条目相匹配。
  • 通过迭代优化与过滤提升摘要质量,再进行最终部署。
  • 将最终摘要集成至RNAcentral数据库,供公众访问。

实验结果

研究问题

  • RQ1大型语言模型能否在极少人工干预的情况下,生成事实准确的ncRNA文献摘要?
  • RQ2在人工评估下,LLM生成摘要的质量与人工整理标准相比如何?
  • RQ3结构化提示与自动化验证能否显著提升LLM生成摘要的事实一致性?
  • RQ4该方法在大规模非编码RNA集合上的可扩展性如何?
  • RQ5此类摘要能否可靠地整合至现有生物知识库(如RNAcentral)?

主要发现

  • 大多数经人工评估的摘要被评为极高质,表明其与专家整理标准高度一致。
  • 该系统成功为超过4,600个非编码RNA生成了摘要,证明了其可扩展性。
  • 自动化检查显著提升了摘要的事实准确性和引用正确性。
  • 思维链提示方法实现了对关键生物学事实(包括功能角色与调控机制)的一致提取。
  • 最终摘要已集成至RNAcentral,使研究人员可公开访问。
  • 结果表明,当结合精心设计的提示与验证流程时,当前的LLM能够生成可靠的基因组学文献摘要。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。