Skip to main content
QUICK REVIEW

[論文レビュー] Automatically applying a credibility appraisal tool to track vaccination-related communications shared on social media

Zubair Shah, Didi Surian|arXiv (Cornell University)|Mar 18, 2019
Misinformation and Its Impacts参考文献 46被引用数 14
ひとこと要約

本研究では、7段階のチェックリストを用いて、Twitterで共有されたワクチン関連のウェブページの信頼性を自動的に評価する機械学習モデルを開発・評価した。最良のモデルは、信頼性の分類において78%の正確性を達成し、リンクされたウェブページの14.4%が低信頼性であることを明らかにした。この低信頼性コンテンツは、最も共有されたリンクを通じて世界中で最大8000万人に達する影響を及ぼした。

ABSTRACT

Background: Tools used to appraise the credibility of health information are time-consuming to apply and require context-specific expertise, limiting their use for quickly identifying and mitigating the spread of misinformation as it emerges. Our aim was to estimate the proportion of vaccination-related posts on Twitter are likely to be misinformation, and how unevenly exposure to misinformation was distributed among Twitter users. Methods: Sampling from 144,878 vaccination-related web pages shared on Twitter between January 2017 and March 2018, we used a seven-point checklist adapted from two validated tools to appraise the credibility of a small subset of 474. These were used to train several classifiers (random forest, support vector machines, and a recurrent neural network with transfer learning), using the text from a web page to predict whether the information satisfies each of the seven criteria. Results: Applying the best performing classifier to the 144,878 web pages, we found that 14.4% of relevant posts to text-based communications were linked to webpages of low credibility and made up 9.2% of all potential vaccination-related exposures. However, the 100 most popular links to misinformation were potentially seen by between 2 million and 80 million Twitter users, and for a substantial sub-population of Twitter users engaging with vaccination-related information, links to misinformation appear to dominate the vaccination-related information to which they were exposed. Conclusions: We proposed a new method for automatically appraising the credibility of webpages based on a combination of validated checklist tools. The results suggest that an automatic credibility appraisal tool can be used to find populations at higher risk of exposure to misinformation or applied proactively to add friction to the sharing of low credibility vaccination information.

研究の動機と目的

  • リアルタイムでの健康フェイクニュース監視において、大規模かつ即時の信頼性評価を手動で行う課題に対処すること。
  • テキストコンテンツに基づいて、ワクチン関連のウェブページの信頼性を予測する自動化された機械学習分類器を開発すること。
  • フォロワーのネットワークデータを用いて、Twitter上で共有された低信頼性ワクチン情報の潜在的影響範囲を推定すること。
  • 低信頼性ワクチンコンテンツが特に多く共有されるサブ集団を特定し、標的的な公衆衛生介入を実施すること。

提案手法

  • 検証済みのツール(DISCERN や QIMR など)を改変した7段階の信頼性チェックリストを用い、474件のワクチン関連ウェブページを手動でラベル付けした。
  • ラベル付けされたウェブページのテキストを用いて、ランダムフォレスト、サポートベクターマシン、トランスファーラーニングを活用した深層学習の再帰的ニューラルネットワークの3つの機械学習モデルを訓練した。
  • 各モデルは、それぞれのウェブページが7つの信頼性基準のいずれを満たしているかを独立して予測するように訓練された。
  • 最も性能の高かったモデルを用い、2017年1月から2018年3月にかけてTwitterで共有された全144,878件のワクチン関連ウェブページを分類した。
  • 潜在的露出数は、各ウェブページを共有したユーザーのフォロワー数の合計によって推定され、視聴者数の上限として用いられた。
  • 低信頼性コンテンツの共有率が特に高いコミュニティを特定するため、サブ集団分析を実施した。

実験結果

リサーチクエスチョン

  • RQ1Twitterで共有されたワクチン関連ウェブページのうち、検証済みの信頼性チェックリストを用いて低信頼性と分類される割合はどの程度か?
  • RQ2テキストコンテンツのみに基づいて、機械学習モデルがワクチン関連ウェブページの信頼性を正確に予測できるか?
  • RQ3低信頼性ワクチンウェブページの潜在的影響範囲はどの程度で、高信頼性コンテンツと比較してどう異なるか?
  • RQ4どのユーザーコミュニティが低信頼性ワクチン情報を特に多く共有しているか?また、それらのコミュニティは孤立しているか、それとも相互に接続されているか?

主な発見

  • 最も性能の高かった分類器は、低信頼性・中信頼性・高信頼性のウェブページを区別する作業において78%の正確性を達成した。
  • モデルは低信頼性ウェブページを96%以上の精度で特定したため、低信頼性コンテンツの同定において高い信頼性を示した。
  • Twitterで共有された144,878件のワクチン関連ウェブページのうち14.4%が低信頼性と分類されたが、それらは合計の潜在的露出数の9.2%にしか満たなかった。
  • 最も人気のある100件の低信頼性リンクは、それぞれ世界中で200万人から8000万人のTwitterユーザーに潜在的に表示された。
  • 全体として低信頼性ウェブページは共有頻度が低かったが、特定のサブ集団内では顕著に多く共有されていた。
  • 本研究は、自動信頼性評価ツールが、フェイクニュースへの露出を低減するためのリアルタイム監視と標的介入を支援できる可能性を示唆している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。