Skip to main content
QUICK REVIEW

[论文解读] OMG U got flu? Analysis of shared health messages for bio-surveillance

Nigel Collier, Nguyễn Trường Sơn|Oct 13, 2011
Misinformation and Its Impacts被引用 13
一句话总结

本文提出利用推特消息中报告的自我保护健康行为(如避免人群或增加卫生习惯)作为疾病暴发的实时生物监测指标。通过使用基于单字词、双字词和正则表达式的监督学习方法,作者将推文分类为四类保护性行为和一类诊断类别,实现了较高的标注者间一致性(kappa = 0.86),并与官方WHO/NREVSS流感数据表现出中等程度的相关性(Spearman等级相关系数),支持社交媒体作为低成本早期预警系统用于疾病暴发监测。

ABSTRACT

Background: Micro-blogging services such as Twitter offer the potential to crowdsource epidemics in real-time. However, Twitter posts ('tweets') are often ambiguous and reactive to media trends. In order to ground user messages in epidemic response we focused on tracking reports of self-protective behaviour such as avoiding public gatherings or increased sanitation as the basis for further risk analysis. Results: We created guidelines for tagging self protective behaviour based on Jones and Salathé (2009)'s behaviour response survey. Applying the guidelines to a corpus of 5283 Twitter messages related to influenza like illness showed a high level of inter-annotator agreement (kappa 0.86). We employed supervised learning using unigrams, bigrams and regular expressions as features with two supervised classifiers (SVM and Naive Bayes) to classify tweets into 4 self-reported protective behaviour categories plus a self-reported diagnosis. In addition to classification performance we report moderately strong Spearman's Rho correlation by comparing classifier output against WHO/NREVSS laboratory data for A(H1N1) in the USA during the 2009-2010 influenza season. Conclusions: The study adds to evidence supporting a high degree of correlation between pre-diagnostic social media signals and diagnostic influenza case data, pointing the way towards low cost sensor networks. We believe that the signals we have modelled may be applicable to a wide range of diseases.

研究动机与目标

  • 开发一种方法,用于在社交媒体中识别自我保护性健康行为,作为疾病暴发的早期指标。
  • 通过聚焦于行为反应而非自我诊断,解决与流感相关的推文存在的模糊性和媒体反应性问题。
  • 建立一个可靠的标注框架,用于对推特数据中的保护性行为进行分类。
  • 评估社交媒体信号与官方流感监测数据之间的相关性。

提出的方法

  • 基于Jones和Salathé(2009)的研究,制定标注指南,用于标注5,283条与流感相关的推文中的自我保护行为。
  • 采用监督学习方法,使用单字词、双字词和正则表达式作为特征,结合支持向量机(SVM)和朴素贝叶斯分类器。
  • 将推文分类为四类保护性行为和一类自我报告的诊断类别。
  • 使用Cohen’s Kappa系数测量标注者间的一致性,获得0.86的高分值。
  • 使用Spearman等级相关系数,将分类器输出与2009–2010年度WHO/NREVSS实验室确诊的A(H1N1)数据进行比较。
  • 利用经过筛选的实时推特消息语料库,将行为信号建模为流行病趋势的代理指标。

实验结果

研究问题

  • RQ1能否通过自然语言处理技术可靠地检测和分类社交媒体中表达的自我保护行为?
  • RQ2这些检测到的行为信号与官方流感监测数据的相关性如何?
  • RQ3在标注社交媒体中的自我保护性健康行为时,人工标注者之间能达到多高的一致性水平?
  • RQ4基于行为反应的社交媒体信号能否作为传统疾病监测系统的低成本替代方案?
  • RQ5社交媒体中保护性行为的趋势在多大程度上能够提前于官方流感病例报告?

主要发现

  • 该标注方案实现了较高的标注者间一致性,Cohen’s Kappa值为0.86,表明在标注保护性行为方面具有很强的可靠性。
  • 监督分类器(SVM和朴素贝叶斯)能够有效区分四类保护性行为与自我报告的诊断类别。
  • 分类器输出与官方WHO/NREVSS A(H1N1)实验室数据之间观察到中等但具有统计学意义的Spearman等级相关系数。
  • 本研究证明,基于行为反应的、尚未确诊的社交媒体信号与现实世界中的疾病趋势存在相关性。
  • 结果支持利用社交媒体作为低成本、实时的传感器网络用于流行病检测的可行性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。