[论文解读] Characterizing the spread of exaggerated news content over social media
本文研究了夸张健康新闻在Twitter上的传播机制,发现关于夸张内容的延迟发布推文更频繁地使用观点和认知类词汇(如'feel'、'realize'),同时使用更少的负面或与死亡相关的词汇。通过语言特征与用户行为分析,作者基于F1得分0.83成功分类出频繁分享夸张内容的用户,揭示了此类用户在行为与语言模式上的显著差异。
In this paper, we consider a dataset comprising press releases about health research from different universities in the UK along with a corresponding set of news articles. First, we do an exploratory analysis to understand how the basic information published in the scientific journals get exaggerated as they are reported in these press releases or news articles. This initial analysis shows that some news agencies exaggerate almost 60\% of the articles they publish in the health domain; more than 50\% of the press releases from certain universities are exaggerated; articles in topics like lifestyle and childhood are heavily exaggerated. Motivated by the above observation we set the central objective of this paper to investigate how exaggerated news spreads over an online social network like Twitter. The LIWC analysis points to a remarkable observation these late tweets are essentially laden in words from opinion and realize categories which indicates that, given sufficient time, the wisdom of the crowd is actually able to tell apart the exaggerated news. As a second step we study the characteristics of the users who never or rarely post exaggerated news content and compare them with those who post exaggerated news content more frequently. We observe that the latter class of users have less retweets or mentions per tweet, have significantly more number of followers, use more slang words, less hyperbolic words and less word contractions. We also observe that the LIWC categories like bio, health, body and negative emotion are more pronounced in the tweets posted by the users in the latter class. As a final step we use these observations as features and automatically classify the two groups achieving an F1 score of 0.83.
研究动机与目标
- 理解来自科研机构新闻稿与新闻报道的夸张健康新闻如何在Twitter等社交媒体平台上传播。
- 识别极少或从不分享夸张内容的用户与频繁分享此类内容的用户在语言与行为上的差异。
- 基于语言与网络特征,构建可预测用户传播夸张新闻倾向的分类模型。
- 研究夸张新闻传播的时间动态,特别是延迟发布推文中语言使用的变化。
提出的方法
- 本研究使用来自英国大学的462条标注新闻稿与668篇新闻文章数据集,标签依据因果主张变化、建议明确性及样本身份变化判断夸张程度。
- 采用语言询问与词频分析(LIWC)对推文进行分析,提取与情绪、认知及社会过程相关的特征,包括'assent'(赞同)、'feel'(感受)、'negative emotion'(负面情绪)与'bio'(生物)等类别。
- 通过用户分享夸张内容的频率对用户进行分类:从未(users_NEX)、偶尔(users_EX1)、两次(users_EX2)或三次及以上(users_EX3)。
- 从用户时间线中提取转发/提及次数、俚语、夸张词汇、缩写形式与推文长度等特征,以比较其发布行为差异。
- 使用上述特征训练监督分类模型,采用分层10折交叉验证与类别权重处理类别不平衡问题。
- 随机森林与XGBoost分类器表现最佳,F1得分为0.83,可有效区分极少或从不分享夸张内容的用户与频繁分享者。
实验结果
研究问题
- RQ1关于夸张健康新闻的推文,其语言内容如何随时间演变,特别是在延迟发布的推文中?
- RQ2哪些行为与语言差异可将频繁分享夸张健康新闻的用户与不分享者区分开?
- RQ3能否利用用户的发帖模式与语言特征预测其传播夸张新闻的可能性?
- RQ4在分享夸张内容与非夸张内容的推文中,LIWC类别如'assent'(赞同)、'anxiety'(焦虑)、'religion'(宗教)与'negative emotion'(负面情绪)有何差异?
主要发现
- 延迟发布的、分享夸张健康新闻的推文在'assent'(赞同)、'feel'(感受)、'religion'(宗教)与'anxiety'(焦虑)等LIWC类别中词汇使用更频繁,表明观点表达与情感认知增强。
- 这些延迟推文中与'death'(死亡)、'sexual'(性)、'sad'(悲伤)及'negative emotion'(负面情绪)相关的词汇使用更少,表明对痛苦或宿命感语言的关注减弱。
- 频繁分享夸张内容的用户拥有显著更多的关注者,使用更多俚语,更少使用夸张词汇,且更少使用缩写形式,相较于极少分享此类内容的用户。
- 语言分析显示,频繁发布夸张内容的用户在其推文中更常使用'bio'(生物)、'health'(健康)、'body'(身体)与'negative emotion'(负面情绪)等LIWC类别的词汇。
- 分类器在区分从未或极少分享夸张内容的用户与频繁分享者时,F1得分为0.83,其中随机森林与XGBoost模型表现最佳。
- 研究结果表明,随着时间推移,群体智慧可能通过语言线索(尤其是观点与认知类语言的增加)识别出夸张内容。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。