Skip to main content
QUICK REVIEW

[论文解读] Nowcasting economic and social data: when and why search engine data fails, an illustration using Google Flu Trends

Paul Ormerod, Rickard Nyman|arXiv (Cornell University)|Aug 1, 2014
Innovation Diffusion and Forecasting参考文献 9被引用 12
一句话总结

本文通过分析搜索行为背后的动机——具体而言,是独立的信息搜寻还是社会影响——来探究谷歌流感趋势为何经常无法准确进行流感实时预测。利用四个国家的搜索数据对巴斯扩散模型进行分析,发现当社会影响占主导时会出现过度预测,而准确预测则与更强的独立搜索动机相关。

ABSTRACT

Obtaining an accurate picture of the current state of the economy is particularly important to central banks and finance ministries, and of epidemics to health ministries. There is increasing interest in the use of search engine data to provide such 'nowcasts' of social and economic indicators. However, people may search for a phrase because they independently want the information, or they may search simply because many others are searching for it. We consider the effect of the motivation for searching on the accuracy of forecasts made using search engine data of contemporaneous social and economic indicators. We illustrate the implications for forecasting accuracy using four episodes in which Google Flu Trends data gave accurate predictions of actual flu cases, and four in which the search data over-predicted considerably. Using a standard statistical methodology, the Bass diffusion model, we show that the independent search for information motive was much stronger in the cases of accurate prediction than in the inaccurate ones. Social influence, the fact that people may search for a phrase simply because many others are, was much stronger in the inaccurate compared to the accurate cases. Search engine data may therefore be an unreliable predictor of contemporaneous indicators when social influence on the decision to search is strong.

研究动机与目标

  • 探究搜索引擎数据(如谷歌流感趋势)为何有时无法准确预测流感病例等实时健康指标。
  • 考察搜索动机(特别是独立信息搜寻与社会影响)在决定基于搜索的实时预测可靠性方面的作用。
  • 评估独立与社会搜索动机的相对强度是否可预测搜索引擎数据在预测同期社会与经济指标方面的准确性。
  • 应用巴斯扩散模型量化不同流感疫情爆发期间独立与社会搜索行为的贡献。
  • 基于行为动机,提供一个识别搜索数据在政策相关实时预测中何时可能可靠的框架。

提出的方法

  • 将巴斯扩散模型应用于谷歌趋势中的流感相关搜索数据,其公式为:$ S(t) = m\frac{(p+q)^2}{p}\frac{e^{-(p+q)t}}{(1+\frac{q}{p}e^{-(p+q)t})^2} $,其中 $ m $ 为累计总搜索量,$ p $ 为独立搜索系数,$ q $ 为受社会影响的搜索系数。
  • 使用 R 语言中的 nlmrt 包进行非线性回归,估算四个国家(美国、瑞士、德国、比利时)在流感高峰期前后搜索数据的 $ p $ 和 $ q $,并比较准确与不准确预测案例。
  • 分析聚焦于从峰值前一个低点到峰值后一个低点的时间段,确保模型拟合的时间窗口一致。
  • 比较准确与过度预测案例中 $ p $ 与 $ q $ 的相对大小,以评估独立与社会搜索动机的主导程度。
  • 使用调整后的 $ R^2 $ 评估模型拟合度,所有案例的值均合理。
  • 研究使用了谷歌流感趋势官网及表 S1 中的具体搜索时间序列数据。

实验结果

研究问题

  • RQ1在何种条件下,搜索引擎数据会无法准确实时预测社会或经济指标?
  • RQ2独立信息搜寻与社会影响如何影响基于搜索的预测可靠性?
  • RQ3独立与社会搜索动机的相对强度在多大程度上可预测谷歌流感趋势预测的准确性?
  • RQ4巴斯扩散模型能否有效区分公共卫生数据中独立与社会搜索行为?
  • RQ5搜索模式中哪些早期信号可表明社会影响占主导,从而削弱预测准确性?

主要发现

  • 在谷歌流感趋势过度预测流感病例的案例中(如美国 2012/13 年、瑞士 2008/09 年),受社会影响的搜索系数 $ q $ 显著大于独立搜索系数 $ p $,表明社会影响强烈。
  • 相反,在预测准确的案例中(如美国 2011/12 年、瑞士 2007/08 年),独立搜索系数 $ p $ 显著大于 $ q $,表明搜索主要由个体信息需求驱动。
  • 巴斯模型拟合的调整后 $ R^2 $ 值在全部八个案例中均合理,证实了该模型对数据的适用性。
  • 研究发现,搜索量迅速上升后缓慢下降的模式表明独立搜索动机较强,而更对称的峰值则表明社会影响占主导。
  • 过度预测始终与社会影响的相对权重更高相关,表明当人们并非为获取信息而搜索,而是因他人在搜索时,搜索数据便不可靠。
  • 结果表明,当个体独立搜索行为占主导时,搜索引擎数据在实时预测中最为有用,而非当社会影响驱动搜索模式时。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。