[論文レビュー] A curated collection of COVID-19 online datasets
本論文は、Twitter、信頼できる保健関連ソース、WHOの世界的な状況報告書から得た、COVID-19関連の1,169件のデータセットを収集したものであり、誤情報の拡散の緩和に関する研究を可能にする。これらのデータセットは、コンテンツ分類、真偽の検証、トーン分析、誤情報の拡散、主題分析、政策感情の追跡を支援し、完全なハイドレーション対応と事実確認システムのベンチマーク化を実現する。
One of the defining moments of the year 2020 is the outbreak of Coronavirus Disease (Covid-19), a deadly virus affecting the body's respiratory system to the point of needing a breathing aid via ventilators. As of June 21, 2020 there are 12,929,306 confirmed cases and 569,738 confirmed deaths across 216 countries, areas or territories. The scale of spread and impact of the pandemic left many nations grappling with preventive and curative approaches. The infamous lockdown measure introduced to mitigate the virus spread has altered many aspects of our social routines in which demand for online-based services skyrocketed. As the virus propagate, so does misinformation and fake news around it via online social media, which seems to favour virality over veracity. With a majority of the populace confined to their homes for a long period, vulnerability to the toxic impact of online misinformation is high. A case in point is the various myths and disinformation associated with the Covid-19, which, if left unchecked, could lead to a catastrophic outcome and hamper the fight against the virus. While the scientific community is actively engaged in identifying the virus treatment, there is a growing interest in combating the associated harmful infodemic. To this end, researchers have been curating and documenting various datasets about Covid-19. In line with existing studies, we provide an expansive collection of curated datasets to support the fight against the pandemic, especially concerning misinformation. The collection consists of 3 categories of Twitter data, information about standard practices from credible sources and a chronicle of global situation reports. We describe how to retrieve the hydrated version of the data and proffer some research problems that could be addressed using the data.
研究の動機と目的
- COVID-19パンデミック期における誤情報やフェイクニュースを研究するための、真のデータやキュレート済みデータセットの不足に対処すること。
- 多様なソースからの構造的でアクセス可能でハイドレート済みのデータを提供することにより、誤情報の拡散ダイナミクスに関する計算科学研究を支援すること。
- ソーシャルメディアにおける誤情報の検出・分類・検証のためのシステムの開発とベンチマーク化を可能にすること。
- パンデミック関連のトピックに関する公的議論の縦断的および主題的分析を可能にすること。これには、政策感情やコミュニティ構造が含まれる。
- 危機的状況下における事実確認および誤情報抑止のための信頼性のあるベンチマークの作成を支援すること。
提案手法
- 主に3つのデータカテゴリの収集:ハッシュタグ、アカウント、時系列フィルタリングを用いたTwitterデータセット、信頼できるソースからの検証済み保健ガイドライン、WHOの世界的な状況報告書。
- Twitterの公式APIを用いて、元のツイートIDを取得し、完全なツイート内容を復元(ハイドレーション)して分析可能にする。
- データの整合性、関連性、真実データとの整合性を確保するためのデータキュレーション技術の適用。
- コンテンツの傾向に基づいて、プロWHOおよびアンチWHOグループにユーザーを分類し、コミュニティ検出と感情分析を可能にする。
- 外部の事実確認リソース(例:IFCN、AFP)を統合し、ツイートの真偽を検証し、ベンチマーク化を支援する。
- ストリーミングAPIを用いてリアルタイムデータを収集し、その後フィルタリングと主題クラスタリングを実施して、パンデミック関連のトピックに焦点を当てる。
実験結果
リサーチクエスチョン
- RQ1ソーシャルメディア上でのパンデミック関連情報の真贋を区別するコンテンツ分類モデルは、どのようにして訓練可能か?
- RQ2権威ある保健関連ソースからの真実データを用いて、ツイートの真偽を検証する最も効果的な方法は何か?
- RQ3誤情報のトーン(例:感情的、否定的、センセーショナル)は、事実情報のトーンとどのように異なるか?
- RQ4パンデミック期における誤情報と正確な情報の拡散における構造的・行動的パターンは何か?
- RQ5ロックダウン政策に対する公的認識の感情は時間経過とともにどのように変化するか?また、感情は政策措置と信頼性を持って関連付けられるか?
主な発見
- 本データセット収集には、Twitter、WHOレポート、検証済み保健関連ソースから1,169件のキュレート済みデータセットが含まれており、誤情報研究の包括的基盤を提供する。
- ハイドレート済みツイートデータの統合により、完全なテキスト分析が可能になり、IDのみのデータセットの制限を克服する。
- 本データセットは、プロWHOおよびアンチWHOグループを含む明確なユーザーコミュニティの特定を可能にし、重複するサブコミュニティの可能性も示す。
- データは縦断的感情分析を可能にし、パンデミック対策に対する公的認識の時間的変化を示す。
- 権威あるソースの統合により、事実確認および誤情報検出システムのベンチマーク化が可能になる。
- 本データセットは、特定のトピックを中心にツイートのフィルタリングとクラスタリングを可能にし、主題分析を促進することで文脈的理解を向上させる。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。