[論文レビュー] Navigating the landscape of COVID-19 research through literature analysis: A bird's eye view
本稿では、13,369編のPubMed論文を対象に名前付きエンティティ認識、クラスタリング、トピックモデリングを用いて、急激に拡大する新型コロナウイルス感染症(COVID-19)関連の学術文献を分析する自然言語処理フレームワークを提示する。本研究では、持続的かつ新規の研究トピックを同定し、主なバイオエンティティ(例:疾患、症状、臓器)をマッピングし、パンデミック期における科学的発見を加速するためのインタラクティブで公開可能な知識ナビゲーションシステムを提供する。
Timely access to accurate scientific literature in the battle with the ongoing COVID-19 pandemic is critical. This unprecedented public health risk has motivated research towards understanding the disease in general, identifying drugs to treat the disease, developing potential vaccines, etc. This has given rise to a rapidly growing body of literature that doubles in number of publications every 20 days as of May 2020. Providing medical professionals with means to quickly analyze the literature and discover growing areas of knowledge is necessary for addressing their question and information needs. In this study we analyze the LitCovid collection, 13,369 COVID-19 related articles found in PubMed as of May 15th, 2020 with the purpose of examining the landscape of literature and presenting it in a format that facilitates information navigation and understanding. We do that by applying state-of-the-art named entity recognition, classification, clustering and other NLP techniques. By applying NER tools, we capture relevant bioentities (such as diseases, internal body organs, etc.) and assess the strength of their relationship with COVID-19 by the extent they are discussed in the corpus. We also collect a variety of symptoms and co-morbidities discussed in reference to COVID-19. Our clustering algorithm identifies topics represented by groups of related terms, and computes clusters corresponding to documents associated with the topic terms. Among the topics we observe several that persist through the duration of multiple weeks and have numerous associated documents, as well several that appear as emerging topics with fewer documents. All the tools and data are publicly available, and this framework can be applied to any literature collection. Taken together, these analyses produce a comprehensive, synthesized view of COVID-19 research to facilitate knowledge discovery from literature.
研究の動機と目的
- 2020年5月時点で、20日ごとに文献数が2倍に増加するなど、急激に拡大するCOVID-19関連の学術文献における情報過多の課題に対処すること。
- 医療従事者や研究者が、進化を続ける研究のあり方を効率的にナビゲートし、臨床意思決定や知識発見を支援できるようにすること。
- NLP技術を用いて、任意の生物医学文献コレクションから主要なエンティティとトピックを抽出・整理するスケーラブルで公開可能なフレームワークの開発。
- 自動クラスタリングと関係性分析を通じて、COVID-19文献における既存の研究テーマと新規の研究テーマを同定すること。
- 高度なテキストマイニングを用いて、SARS-CoV-2と関連する合併症、症状、生物学的エンティティを発見すること。
提案手法
- 13,369編のCOVID-19関連PubMed論文から、最新の名前付きエンティティ認識(NER)ツールを用いてバイオエンティティ(例:疾患、臓器、遺伝子)を抽出する。
- コーパス内での共起頻度と文脈的関連性に基づいて、バイオエンティティとCOVID-19との間の関係の強さを測定する。
- クラスタリングアルゴリズムを用いて関連語をグループ化し、一貫性のある研究トピック(持続的・新規のものも含む)を同定する。
- 複数週間にわたるトピッククラスタの進化を追跡することで、時間的トレンドを分析し、新規の研究分野を検出する。
- 分類および情報検索技術を統合して、文献を構造化・要約し、ナビゲーション性を向上させる。
- すべてのツール、データ、結果を公開し、今後の文献分析タスクにおける再現性と再利用を支援する。
実験結果
リサーチクエスチョン
- RQ1科学的文献に反映されたCOVID-19の文脈において、最も頻繁に議論されている症状と合併症は何か?
- RQ2SARS-CoV-2と強く関連付けられている生物学的エンティティ(例:臓器、タンパク質)は何か?
- RQ3COVID-19関連の研究トピックの中で、時間経過にかかわらず持続的であるものと、初期段階で文書数が少ないが成長著しいものとは何か?
- RQ4NLP技術を大規模な生物医学文献に効果的に適用することで、知識を抽出・整理し、迅速な発見を可能にする方法は何か?
- RQ5自動化された文献分析は、COVID-19パンデミックのようなグローバルな健康危機への的確な理解と迅速な対応をどの程度支援できるか?
主な発見
- 本研究では、文書数が多く、研究の核心的分野にわたり持続的な科学的関心が集まっている複数の恒久的な研究トピックが同定された。
- 初期段階の文書数が少ないが、成長著しい新規トピックが複数検出され、新規研究分野の兆しを示した。
- 症状や合併症が体系的に抽出され、SARS-CoV-2と結びつけられ、臨床的関連性の構造的概要が得られた。
- 肺、サイトコイニン、ACE2受容体といった主要なバイオエンティティが顕著に議論されており、それらの生物学的関連性が裏付けられた。
- 本フレームワークは、疾患、臓器、症状の間の関係を的確にマッピングし、文献全体の統合的視覚化を可能にした。
- すべてのツールとデータが公開されており、他の疾患分野や文献コレクションへの応用・拡張が可能である。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。