Skip to main content
QUICK REVIEW

[论文解读] Social media emotion macroscopes reflect emotional experiences in society at large

David García, Max Pellert|arXiv (Cornell University)|Jul 28, 2021
Mental Health Research Topics参考文献 5被引用 16
一句话总结

本研究通过每周一次的英国代表性调查,验证了基于性别重标度文本分析的社交媒体情绪显微镜(即聚合的Twitter情绪指标)的有效性,结果显示Twitter情绪时间线与自我报告情绪之间存在强烈相关性(r = 0.69–0.81)。研究发现,社交媒体不仅反映个体情绪,还能捕捉集体情绪状态,主要通过第三人称指代反映对他人情绪的感知。

ABSTRACT

Social media generate data on human behaviour at large scales and over long periods of time, posing a complementary approach to traditional methods in the social sciences. Millions of texts from social media can be processed with computational methods to study emotions over time and across regions. However, recent research has shown weak correlations between social media emotions and affect questionnaires at the individual level and between static regional aggregates of social media emotion and subjective well-being at the population level, questioning the validity of social media data to study emotions. Yet, to date, no research has tested the validity of social media emotion macroscopes to track the temporal evolution of emotions at the level of a whole society. Here we present a pre-registered prediction study that shows how gender-rescaled time series of Twitter emotional expression at the national level substantially correlate with aggregates of self-reported emotions in a weekly representative survey in the United Kingdom. A follow-up exploratory analysis shows a high prevalence of third-person references in emotionally-charged tweets, indicating that social media data provide a way of social sensing the emotions of others rather than just the emotional experiences of users. These results show that, despite the issues that social media have in terms of representativeness and algorithmic confounding, the combination of advanced text analysis methods with user demographic information in social media emotion macroscopes can provide measures that are informative of the general population beyond social media users.

研究动机与目标

  • 检验大规模Twitter情绪聚合是否能可靠反映社会层面的国家情绪趋势。
  • 通过在情绪指标中应用性别重标度,解决社交媒体数据中代表性不足和性别偏差的问题。
  • 探究社交媒体是否不仅捕捉个人情绪,还通过第三人称指代反映对他人情绪的感知。
  • 通过预注册的研究设计,验证情绪显微镜在时间维度上的预测能力。
  • 比较基于词典的方法与基于深度学习的情绪检测方法在追踪群体情感状态方面的表现。

提出的方法

  • 一项预注册的预测研究,将基于Twitter的情绪时间线与2019年6月至2021年6月期间英国YouGov的每周调查数据进行对比。
  • 使用基于LIWC的词典方法和微调后的RoBERTa分类器,在15亿条地理位置标注的英国Twitter推文中检测12种情绪。
  • 应用性别重标度以纠正Twitter上男性用户过度代表的问题,提升情绪指标的人口代表性。
  • 追踪情绪化推文中第三人称代词(如they, him, she)的使用情况,以评估对他人情绪的社会感知。
  • 采用卡方检验比较情绪相关与非情绪相关推文中代词频率的差异,识别出显著差异。
  • 报告了调查数据与Twitter情绪时间线之间相关系数的95%置信区间,确保统计稳健性。

实验结果

研究问题

  • RQ1经过性别重标度的Twitter情绪时间线能否预测代表性调查所测量的国家层面情绪趋势?
  • RQ2Twitter情绪检测与调查报告情绪之间的相关性在不同性别和情绪类型之间如何变化?
  • RQ3情绪化推文中包含第三人称指代的程度如何,表明对他人情绪的社会感知?
  • RQ4最先进的NLP模型(如RoBERTa)在捕捉群体情绪方面是否优于传统词典方法?
  • RQ5使用Twitter数据的预注册预测模型是否具有时间上的泛化能力,避免事后数据拟合?

主要发现

  • 经过性别重标度的Twitter情绪时间线与英国调查中自我报告的悲伤情绪表现出显著相关性(r = 0.698),证实其作为宏观情绪传感器的有效性。
  • 对于焦虑情绪,女性的相关性达到r = 0.811,男性为r = 0.724,表明与调查数据高度一致。
  • 第三人称代词在含情绪词的推文中显著更常见——例如,悲伤相关推文中占41.94%,而非悲伤推文中为36.07%,表明存在对他人情绪的社会感知。
  • 使用性别重标度显著提升了代表性,降低了因Twitter用户性别构成中男性占主导而产生的偏差。
  • 在大多数情绪类别中,基于词典的方法(LIWC)比RoBERTa获得更高的相关性,尤其在焦虑和挫败感方面表现更优。
  • 即使未进行性别重标度,'害怕'和'挫败'等情绪也表现出强相关性(r = 0.794 和 r = 0.649),表明其信号检测具有鲁棒性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。