[论文解读] Automatically applying a credibility appraisal tool to track vaccination-related communications shared on social media
本研究开发并评估了机器学习模型,以自动评估在推特上分享的疫苗相关网页的可信度,采用七分制检查清单。最佳模型在分类可信度方面达到了78%的准确率,并发现14.4%的链接网页可信度较低,通过传播最广的链接,全球影响范围高达8000万用户。
Background: Tools used to appraise the credibility of health information are time-consuming to apply and require context-specific expertise, limiting their use for quickly identifying and mitigating the spread of misinformation as it emerges. Our aim was to estimate the proportion of vaccination-related posts on Twitter are likely to be misinformation, and how unevenly exposure to misinformation was distributed among Twitter users. Methods: Sampling from 144,878 vaccination-related web pages shared on Twitter between January 2017 and March 2018, we used a seven-point checklist adapted from two validated tools to appraise the credibility of a small subset of 474. These were used to train several classifiers (random forest, support vector machines, and a recurrent neural network with transfer learning), using the text from a web page to predict whether the information satisfies each of the seven criteria. Results: Applying the best performing classifier to the 144,878 web pages, we found that 14.4% of relevant posts to text-based communications were linked to webpages of low credibility and made up 9.2% of all potential vaccination-related exposures. However, the 100 most popular links to misinformation were potentially seen by between 2 million and 80 million Twitter users, and for a substantial sub-population of Twitter users engaging with vaccination-related information, links to misinformation appear to dominate the vaccination-related information to which they were exposed. Conclusions: We proposed a new method for automatically appraising the credibility of webpages based on a combination of validated checklist tools. The results suggest that an automatic credibility appraisal tool can be used to find populations at higher risk of exposure to misinformation or applied proactively to add friction to the sharing of low credibility vaccination information.
研究动机与目标
- 为解决在实时健康虚假信息监测中大规模、实时手动评估可信度的挑战。
- 开发自动化的机器学习分类器,基于文本内容预测疫苗相关网页的可信度。
- 利用用户关注网络数据,估算在推特上分享的低可信度疫苗信息的潜在传播范围。
- 识别低可信度疫苗内容在特定子群体中被不成比例分享的情况,以支持有针对性的公共卫生干预。
提出的方法
- 采用经验证的工具(如DISCERN和QIMR)改编的七分制可信度检查清单,对474个疫苗相关网页进行人工标注。
- 使用标注网页的文本训练三种机器学习模型:随机森林、支持向量机以及采用迁移学习的深度学习循环神经网络。
- 模型被训练以独立预测每个网页是否满足七项可信度标准中的每一项。
- 将表现最佳的模型应用于分类2017年1月至2018年3月期间在推特上分享的全部144,878个疫苗相关网页。
- 通过累加分享每个网页的用户的总关注者数来估算潜在传播范围,作为受众覆盖范围的上限。
- 开展子群体分析,以识别低可信度内容传播率较高的社区。
实验结果
研究问题
- RQ1使用经验证的可信度检查清单,推特上分享的疫苗相关网页中,有多少比例被归类为可信度较低?
- RQ2仅基于文本内容,机器学习模型能否准确预测疫苗相关网页的可信度?
- RQ3低可信度疫苗网页在推特上的潜在传播范围是多少?与高可信度内容相比如何?
- RQ4哪些用户群体最可能分享低可信度的疫苗信息?这些群体是更孤立还是更互联?
主要发现
- 表现最佳的分类器在区分低、中、高可信度网页方面达到了78%的准确率。
- 该模型对低可信度网页的精确度超过96%,表明在识别低可信度内容方面具有高度可靠性。
- 在144,878个在推特上分享的疫苗相关网页中,14.4%被归类为可信度较低,但仅占总潜在传播量的9.2%。
- 传播最广的100个低可信度链接,每个在全球范围内潜在影响的用户数在200万至8000万之间。
- 低可信度网页整体分享频率较低,但在特定子群体中被不成比例地传播。
- 本研究表明,自动化可信度评估工具可支持实时监测,并助力有针对性的干预措施,以减少用户接触虚假信息的风险。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。