Skip to main content
QUICK REVIEW

[论文解读] No Love Among Haters: Negative Interactions Reduce Hate Community Engagement

Daniel Hickey, Matheus Schmitz|arXiv (Cornell University)|Mar 23, 2023
Hate Speech and Cyberbullying Detection被引用 4
一句话总结

本研究运用因果推断表明,负面互动——尤其是具有毒性、攻击性和敌意的回复——会降低新用户在仇恨型Reddit社区中的参与度,尽管这些社区存在监管不足,但这一现象出人意料地限制了招募。通过使用Perspective API和VADER,研究发现此类敌意会阻碍参与行为,模拟结果进一步证实,若回复更加友善,仇恨型子版块的参与度将显著提升,而非仇恨型社区则保持稳定。

ABSTRACT

While online hate groups pose significant risks to the health of online platforms and safety of marginalized groups, little is known about what causes users to become active in hate groups and the effect of social interactions on furthering their engagement. We address this gap by first developing tools to find hate communities within Reddit, and then augment 11 subreddits extracted with 14 known hateful subreddits (25 in total). Using causal inference methods, we evaluate the effect of replies on engagement in hateful subreddits by comparing users who receive replies to their first comment (the treatment) to equivalent control users who do not. We find users who receive replies are less likely to become engaged in hateful subreddits than users who do not, while the opposite effect is observed for a matched sample of similar-sized non-hateful subreddits. Using the Google Perspective API and VADER, we discover that hateful community first-repliers are more toxic, negative, and attack the posters more often than non-hateful first-repliers. In addition, we uncover a negative correlation between engagement and attacks or toxicity of first-repliers. We simulate the cumulative engagement of hateful and non-hateful subreddits under the contra-positive scenario of friendly first-replies, finding that attacks dramatically reduce engagement in hateful subreddits. These results counter-intuitively imply that, although under-moderated communities allow hate to fester, the resulting environment is such that direct social interaction does not encourage further participation, thus endogenously constraining the harmful role that these communities could play as recruitment venues for antisocial beliefs.

研究动机与目标

  • 理解在线仇恨社区中用户参与度的驱动因素,特别是社交互动的作用。
  • 探究初始互动——尤其是回复——是否影响新用户在仇恨型子版块中持续参与的可能性。
  • 比较仇恨型子版块与非仇恨型社区中负面互动的影响差异。
  • 量化回复中的毒性、攻击性与负面情绪与用户参与度下降之间的相关性。
  • 模拟更友善的首次回复对仇恨社区长期参与度的影响。

提出的方法

  • 开发了一种新方法,从Reddit数据中检测仇恨型子版块,使用已知的仇恨型子版块作为训练集。
  • 通过对比收到其首次帖子回复的用户(处理组)与未收到回复的相似用户(对照组)实施因果推断。
  • 使用Google的Perspective API和VADER分析回复中的毒性、对发帖者的攻击性以及负面情绪。
  • 构建模拟模型,估算在‘反向积极’情境下——即首次回复为非毒性且非攻击性——的累积参与度。
  • 通过比较规模和活跃度相近的子版块,控制预先存在的仇恨言论水平。
  • 使用人工标注数据验证Perspective API指标,获得AUC值分别为0.85(对发帖者的攻击)和0.86(毒性),表明其可靠性。
Figure 1: Schematic of hateful subreddit growth simulation. A mixed effect logistic regression model predicts whether a user continues posting or leaves the subreddit. Predictions are counted to calculate the cumulative number of engaged users in a subreddit.
Figure 1: Schematic of hateful subreddit growth simulation. A mixed effect logistic regression model predicts whether a user continues posting or leaves the subreddit. Predictions are counted to calculate the cumulative number of engaged users in a subreddit.

实验结果

研究问题

  • RQ1收到对其首次帖子的回复,会增加还是减少用户在仇恨型子版块中持续参与的可能性?
  • RQ2回复的情感与语言特征(毒性、负面情绪、攻击性)如何影响仇恨社区中的用户留存?
  • RQ3仇恨型子版块中回复的影响,与非仇恨型子版块中的影响相比如何?
  • RQ4如果仇恨型子版块的首次回复不具敌意,对参与度的累积影响将如何?
  • RQ5所使用的语言模型(Perspective API、VADER)在检测仇恨社区话语中的敌意方面是否可靠?

主要发现

  • 收到首次帖子回复的用户,在仇恨型子版块中持续活跃的可能性显著降低,这与非仇恨型社区的趋势相反。
  • 即使在控制了仇恨言论使用水平后,仇恨型子版块的首次回复仍显著比非仇恨型子版块更具毒性、负面情绪更强且更具攻击性。
  • 首次回复的毒性、对发帖者的攻击性以及负面情绪程度与用户在仇恨型子版块中参与度下降之间存在强烈负相关性。
  • 模拟结果显示,若将敌意回复替换为中性或友好回复,仇恨型子版块的长期参与度将显著提升,表明敌意行为本身具有自我限制机制。
  • 在相同模拟条件下,非仇恨型子版块的参与度变化极小,因为其本身毒性与负面情绪水平已很低。
  • Perspective API的‘对发帖者的攻击’与‘毒性’指标AUC值分别为0.85和0.86,证实其在本情境下的可靠性。
Figure 2: Replies to comments in hateful subreddits lead to significantly less engagement than replies to comments in non-hateful subreddits. Distributions of engagement risk ratios for different subreddit types, separated by users who make comments as their first post (A) and users who make submiss
Figure 2: Replies to comments in hateful subreddits lead to significantly less engagement than replies to comments in non-hateful subreddits. Distributions of engagement risk ratios for different subreddit types, separated by users who make comments as their first post (A) and users who make submiss

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。