Skip to main content
QUICK REVIEW

[Paper Review] No Love Among Haters: Negative Interactions Reduce Hate Community Engagement

Daniel Hickey, Matheus Schmitz|arXiv (Cornell University)|Mar 23, 2023
Hate Speech and Cyberbullying Detection4 citations
TL;DR

This study uses causal inference to show that negative interactions—particularly toxic, attacking, and hostile replies—reduce new users' engagement in hateful Reddit communities, counterintuitively limiting recruitment despite under-moderation. Using Perspective API and VADER, it finds that such hostility deters participation, and simulations confirm that friendlier replies would significantly increase engagement in hate subreddits, while non-hateful communities remain stable.

ABSTRACT

While online hate groups pose significant risks to the health of online platforms and safety of marginalized groups, little is known about what causes users to become active in hate groups and the effect of social interactions on furthering their engagement. We address this gap by first developing tools to find hate communities within Reddit, and then augment 11 subreddits extracted with 14 known hateful subreddits (25 in total). Using causal inference methods, we evaluate the effect of replies on engagement in hateful subreddits by comparing users who receive replies to their first comment (the treatment) to equivalent control users who do not. We find users who receive replies are less likely to become engaged in hateful subreddits than users who do not, while the opposite effect is observed for a matched sample of similar-sized non-hateful subreddits. Using the Google Perspective API and VADER, we discover that hateful community first-repliers are more toxic, negative, and attack the posters more often than non-hateful first-repliers. In addition, we uncover a negative correlation between engagement and attacks or toxicity of first-repliers. We simulate the cumulative engagement of hateful and non-hateful subreddits under the contra-positive scenario of friendly first-replies, finding that attacks dramatically reduce engagement in hateful subreddits. These results counter-intuitively imply that, although under-moderated communities allow hate to fester, the resulting environment is such that direct social interaction does not encourage further participation, thus endogenously constraining the harmful role that these communities could play as recruitment venues for antisocial beliefs.

Motivation & Objective

  • To understand what drives user engagement in online hate communities, particularly the role of social interactions.
  • To investigate whether initial interactions—especially replies—affect a newcomer’s likelihood of continued participation in hateful subreddits.
  • To compare the impact of negative interactions in hateful subreddits versus non-hateful communities.
  • To quantify how toxicity, attacks, and negativity in replies correlate with reduced user engagement.
  • To simulate the effect of friendlier first-replies on long-term engagement in hate communities.

Proposed method

  • Developed a novel method to detect hateful subreddits from Reddit data, using known hate subreddits as a training set.
  • Applied causal inference by comparing users who received replies to their first post (treatment) with similar users who did not (control).
  • Used Google’s Perspective API and VADER to analyze replies for toxicity, attack on commenter, and negative sentiment.
  • Constructed a simulation model to estimate cumulative engagement under a 'contra-positive' scenario where first-replies are non-toxic and non-attacking.
  • Controlled for pre-existing hate speech levels by comparing subreddits of similar size and activity.
  • Validated Perspective API metrics using human-labeled data, achieving AUC scores of 0.85 (attack on commenter) and 0.86 (toxicity).
Figure 1: Schematic of hateful subreddit growth simulation. A mixed effect logistic regression model predicts whether a user continues posting or leaves the subreddit. Predictions are counted to calculate the cumulative number of engaged users in a subreddit.
Figure 1: Schematic of hateful subreddit growth simulation. A mixed effect logistic regression model predicts whether a user continues posting or leaves the subreddit. Predictions are counted to calculate the cumulative number of engaged users in a subreddit.

Experimental results

Research questions

  • RQ1Does receiving a reply to a first post increase or decrease the likelihood of continued engagement in hateful subreddits?
  • RQ2How do the emotional and linguistic characteristics of replies (toxicity, negativity, attacks) affect user retention in hate communities?
  • RQ3How does the impact of replies in hateful subreddits compare to that in non-hateful subreddits?
  • RQ4What would be the cumulative effect on engagement if first-replies in hateful subreddits were less hostile?
  • RQ5Are the language models used (Perspective API, VADER) reliable for detecting antagonism in hate community discourse?

Key findings

  • Users who received replies to their first post were significantly less likely to remain active in hateful subreddits, contrary to the trend in non-hateful communities.
  • First-replies in hateful subreddits were substantially more toxic, negative, and attacking than those in non-hateful subreddits, even after controlling for hate speech usage.
  • There was a strong negative correlation between the toxicity, attack on commenter, and negativity of first-replies and the probability of user engagement in hateful subreddits.
  • Simulations showed that replacing hostile replies with neutral or friendly ones would dramatically increase long-term engagement in hateful subreddits, suggesting that hostility acts as a self-limiting mechanism.
  • Non-hateful subreddits showed minimal change in engagement under the same simulation, as they already had low levels of toxicity and negativity.
  • The Perspective API’s attack on commenter and toxicity metrics achieved AUC scores of 0.85 and 0.86, respectively, confirming their reliability for this context.
Figure 2: Replies to comments in hateful subreddits lead to significantly less engagement than replies to comments in non-hateful subreddits. Distributions of engagement risk ratios for different subreddit types, separated by users who make comments as their first post (A) and users who make submiss
Figure 2: Replies to comments in hateful subreddits lead to significantly less engagement than replies to comments in non-hateful subreddits. Distributions of engagement risk ratios for different subreddit types, separated by users who make comments as their first post (A) and users who make submiss

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.