Skip to main content
QUICK REVIEW

[论文解读] COVIDScholar: An automated COVID-19 research aggregation and analysis platform

Amalie Trewartha, John Dagdelen|arXiv (Cornell University)|Dec 7, 2020
COVID-19 diagnosis using AI被引用 6
一句话总结

COVIDScholar 是一个基于 NLP 的自动化平台,通过可扩展的数据管道和搜索界面,聚合并分析来自多个学科(包括生物医学、物理科学、社会科学和人文学科)的超过 81,000 篇与 COVID-19 相关的科学文献。该平台使研究人员能够从快速增长的文献中发现可操作的见解,其门户每周服务超过 2,000 名独立用户,并揭示出研究产出在 2020 年 4 月达到峰值,且心理健康和封锁相关研究显著增加。

ABSTRACT

The ongoing COVID-19 pandemic has had far-reaching effects throughout society, and science is no exception. The scale, speed, and breadth of the scientific community's COVID-19 response has lead to the emergence of new research literature on a remarkable scale -- as of October 2020, over 81,000 COVID-19 related scientific papers have been released, at a rate of over 250 per day. This has created a challenge to traditional methods of engagement with the research literature; the volume of new research is far beyond the ability of any human to read, and the urgency of response has lead to an increasingly prominent role for pre-print servers and a diffusion of relevant research across sources. These factors have created a need for new tools to change the way scientific literature is disseminated. COVIDScholar is a knowledge portal designed with the unique needs of the COVID-19 research community in mind, utilizing NLP to aid researchers in synthesizing the information spread across thousands of emergent research articles, patents, and clinical trials into actionable insights and new knowledge. The search interface for this corpus, https://covidscholar.org, now serves over 2000 unique users weekly. We present also an analysis of trends in COVID-19 research over the course of 2020.

研究动机与目标

  • 应对迅速扩大的 COVID-19 科学文献所带来的信息过载挑战,该挑战已超出人工审阅的能力范围。
  • 通过创建统一、可搜索的文献集合,克服研究分散在预印本服务器、期刊和存储库中的问题。
  • 通过基于 NLP 的分析,使研究人员能够从物理科学、公共卫生和社会科学等多样化领域中提取可操作的见解。
  • 开发一种可扩展的、领域无关的科学文献聚合基础设施,可超越疫情背景进行应用。

提出的方法

  • 自动从 14 个开放获取和预印本来源(包括 arXiv、bioRxiv、medRxiv、CORD-19 和 LitCovid)抓取并摄入新发表的文献、专利和临床试验。
  • 使用 DOI、PubMed ID 和基于标题的匹配方法,通过去重管道将同一文献的多个版本合并为单一、优先排序的记录。
  • 使用 NLP 模型对文献进行与 COVID-19 的相关性、主题、学科和领域分类,分类系统基于语料库训练而成的分层分类体系。
  • 使用 pdfminer 和 OCR 从 PDF 中解析全文内容,但受限于文本质量,仅支持文本搜索功能,无法用于 NLP 模型。
  • 从 Crossref 和 OpenCitations 获取元数据和引用数据,构建统一、结构化的语料库。
  • 在 https://covidscholar.org 部署基于 Web 的搜索门户,索引语料库并支持研究人员进行语义和关键词查询。

实验结果

研究问题

  • RQ1从 2020 年 1 月到 12 月,COVID-19 研究文献的体量和跨学科分布如何演变?
  • RQ2在整个疫情期间,主导研究主题和关注重点的变化趋势是什么,特别是在社会科学和心理健康领域?
  • RQ3科学界的产出在多大程度上达到了饱和点,这一现象何时出现?
  • RQ4重大政策事件(如美国的经济刺激法案和世卫组织的疫情宣布)与研究产出变化之间是否存在相关性?
  • RQ5NLP 技术能否在疫情背景下有效识别并揭示非结构化科学文献中的潜在知识?

主要发现

  • 截至 2020 年 10 月,从 14 个来源收集了超过 81,000 篇与 COVID-19 相关的科学文献,2020 年 5 月的月度产出峰值约为 8,000 篇。
  • 研究产出在 2020 年 4 月已达到饱和,尽管 3 月美国通过了经济刺激法案等重大政策事件,此后未观察到产出的显著增长。
  • 物理与医学科学领域的论文占比从 1 月的 5% 和 15% 分别上升至 4 月的 8% 和 20%,表明研究重点发生转移。
  • 关于心理健康和封锁相关主题的文献比例从 3 月的 0% 上升至 4 月至 6 月的 6%–8%,与全球封锁措施同步。
  • 人文学科和社会科学在研究中所占份额持续增长,尤其在心理影响方面,这与广泛实施社会隔离措施的开始密切相关。
  • 该平台每周服务超过 2,000 名独立用户,表明其在科学界中具有强大的采用率和实用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。