Skip to main content
QUICK REVIEW

[论文解读] Navigating the landscape of COVID-19 research through literature analysis: A bird's eye view

Lana Yeganova, Rezarta Islamaj|arXiv (Cornell University)|Aug 7, 2020
Artificial Intelligence in Healthcare and Education参考文献 10被引用 6
一句话总结

本文提出了一种自然语言处理框架,用于分析迅速扩展的新冠肺炎文献,对13,369篇PubMed文章进行命名实体识别、聚类和主题建模。该框架识别出持续存在和新兴的研究主题,映射关键生物实体(如疾病、症状、器官),并提供一个交互式、公开可用的知识导航系统,以加速大流行期间的科学发现。

ABSTRACT

Timely access to accurate scientific literature in the battle with the ongoing COVID-19 pandemic is critical. This unprecedented public health risk has motivated research towards understanding the disease in general, identifying drugs to treat the disease, developing potential vaccines, etc. This has given rise to a rapidly growing body of literature that doubles in number of publications every 20 days as of May 2020. Providing medical professionals with means to quickly analyze the literature and discover growing areas of knowledge is necessary for addressing their question and information needs. In this study we analyze the LitCovid collection, 13,369 COVID-19 related articles found in PubMed as of May 15th, 2020 with the purpose of examining the landscape of literature and presenting it in a format that facilitates information navigation and understanding. We do that by applying state-of-the-art named entity recognition, classification, clustering and other NLP techniques. By applying NER tools, we capture relevant bioentities (such as diseases, internal body organs, etc.) and assess the strength of their relationship with COVID-19 by the extent they are discussed in the corpus. We also collect a variety of symptoms and co-morbidities discussed in reference to COVID-19. Our clustering algorithm identifies topics represented by groups of related terms, and computes clusters corresponding to documents associated with the topic terms. Among the topics we observe several that persist through the duration of multiple weeks and have numerous associated documents, as well several that appear as emerging topics with fewer documents. All the tools and data are publicly available, and this framework can be applied to any literature collection. Taken together, these analyses produce a comprehensive, synthesized view of COVID-19 research to facilitate knowledge discovery from literature.

研究动机与目标

  • 解决截至2020年5月,新冠肺炎科学文献快速增长(每20天翻倍)带来的信息过载挑战。
  • 使医疗专业人员和研究人员能够高效地导航并理解不断演变的研究格局,以支持临床决策和知识发现。
  • 开发一种可扩展的、公开可用的框架,将NLP技术应用于任何生物医学文献集合,以提取和组织关键实体与主题。
  • 通过自动聚类和关系分析,识别新冠肺炎文献中既有的和新兴的研究主题。
  • 通过高级文本挖掘技术,促进与SARS-CoV-2相关的共病、症状和生物实体的发现。

提出的方法

  • 应用最先进的命名实体识别(NER)工具,从13,369篇与新冠肺炎相关的PubMed文章中提取生物实体(如疾病、器官、基因)。
  • 基于语料库中共同出现频率和上下文相关性,衡量生物实体与新冠肺炎之间关系的强度。
  • 使用聚类算法对相关术语进行分组,识别出不同的研究主题,包括持续存在和新兴的主题。
  • 通过跟踪多个星期内主题聚类的演变,分析时间趋势,以检测新兴的研究领域。
  • 整合分类和信息检索技术,对文献进行结构化和摘要化处理,以改善导航体验。
  • 将所有工具、数据和结果公开发布,以支持未来文献分析任务的可重现性和再利用。

实验结果

研究问题

  • RQ1在科学文献中,新冠肺炎背景下最常讨论的症状和共病是什么?
  • RQ2在已发表的研究中,哪些生物实体(如器官、蛋白质)与SARS-CoV-2关联最强?
  • RQ3与新冠肺炎相关的哪些研究主题随时间持续存在,哪些是初期文档较少但正在兴起的新兴主题?
  • RQ4NLP技术如何有效应用于大规模生物医学文献,以提取和组织知识以实现快速发现?
  • RQ5自动化文献分析在多大程度上可以支持对新冠肺炎大流行等全球健康危机的及时理解与应对?

主要发现

  • 研究识别出多个文档数量高的持久性研究主题,表明核心新冠肺炎研究领域持续受到科学界的关注。
  • 检测到若干初期文档较少的新兴主题,表明这些是初现但正在增长的研究方向。
  • 系统性地提取并关联了多种症状和共病与SARS-CoV-2,提供了临床关联的结构化概览。
  • 关键生物实体如肺部、细胞因子和ACE2受体被显著频繁提及,反映出其生物学相关性。
  • 该框架成功映射了疾病、器官和症状之间的关系,实现了对文献的综合视图。
  • 所有工具和数据均可公开访问,支持在其他疾病背景或文献集合中的再利用与扩展。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。