Skip to main content
QUICK REVIEW

[論文レビュー] No Love Among Haters: Negative Interactions Reduce Hate Community Engagement

Daniel Hickey, Matheus Schmitz|arXiv (Cornell University)|Mar 23, 2023
Hate Speech and Cyberbullying Detection被引用数 4
ひとこと要約

本研究では因果推論を用いて、特に有毒で攻撃的・敵対的な返信が、憎悪を示すRedditコミュニティにおける新規ユーザーの参加意欲を低下させることを示した。これは、過小管理が続く中で予想に反して新規参加者の減少を引き起こす。Perspective APIとVADERを用いた分析から、こうした敵対的反応が参加を妨げることを特定し、シミュレーションにより、より親しみやすい返信が憎悪系サブレdditでの参加を著しく増加させることを確認した。一方、非憎悪的コミュニティでは参加が安定していることが分かった。

ABSTRACT

While online hate groups pose significant risks to the health of online platforms and safety of marginalized groups, little is known about what causes users to become active in hate groups and the effect of social interactions on furthering their engagement. We address this gap by first developing tools to find hate communities within Reddit, and then augment 11 subreddits extracted with 14 known hateful subreddits (25 in total). Using causal inference methods, we evaluate the effect of replies on engagement in hateful subreddits by comparing users who receive replies to their first comment (the treatment) to equivalent control users who do not. We find users who receive replies are less likely to become engaged in hateful subreddits than users who do not, while the opposite effect is observed for a matched sample of similar-sized non-hateful subreddits. Using the Google Perspective API and VADER, we discover that hateful community first-repliers are more toxic, negative, and attack the posters more often than non-hateful first-repliers. In addition, we uncover a negative correlation between engagement and attacks or toxicity of first-repliers. We simulate the cumulative engagement of hateful and non-hateful subreddits under the contra-positive scenario of friendly first-replies, finding that attacks dramatically reduce engagement in hateful subreddits. These results counter-intuitively imply that, although under-moderated communities allow hate to fester, the resulting environment is such that direct social interaction does not encourage further participation, thus endogenously constraining the harmful role that these communities could play as recruitment venues for antisocial beliefs.

研究の動機と目的

  • オンラインの憎悪コミュニティにおけるユーザー参加の要因を理解すること、特に社会的相互作用の役割を明らかにすること。
  • 初期の相互作用、特に返信が新規参加者が憎悪系サブレdditに継続的に参加する可能性に与える影響を調査すること。
  • 憎悪系サブレdditと非憎悪コミュニティにおける否定的相互作用の影響を比較すること。
  • 返信における攻撃性、攻撃的発言、否定的感情の度合いが、ユーザー参加意欲に与える相関関係を定量化すること。
  • 友好的な最初の返信が、憎悪コミュニティにおける長期的参加に与える影響をシミュレーションで推定すること。

提案手法

  • 既知の憎悪系サブレdditをトレーニングデータとして用い、Redditデータから憎悪系サブレdditを同定する新規手法を開発した。
  • 処置群(最初の投稿に対して返信を受けたユーザー)と対照群(返信を受けなかった類似ユーザー)を比較することで、因果推論を適用した。
  • GoogleのPerspective APIとVADERを用い、返信の有毒性、コメント投稿者への攻撃、否定的感情を分析した。
  • 最初の返信が有毒でも攻撃的でもない状況を想定した「逆方向の肯定的」シナリオを想定し、累積参加度の推定を行うシミュレーションモデルを構築した。
  • 同規模・同活動レベルのサブレddit同士を比較することで、事前の憎悪スピーチのレベルに影響されないよう制御した。
  • 人間によるラベル付けデータを用いてPerspective APIの指標を検証し、攻撃的発言のAUCスコアが0.85、有毒性のAUCスコアが0.86を達成した。
Figure 1: Schematic of hateful subreddit growth simulation. A mixed effect logistic regression model predicts whether a user continues posting or leaves the subreddit. Predictions are counted to calculate the cumulative number of engaged users in a subreddit.
Figure 1: Schematic of hateful subreddit growth simulation. A mixed effect logistic regression model predicts whether a user continues posting or leaves the subreddit. Predictions are counted to calculate the cumulative number of engaged users in a subreddit.

実験結果

リサーチクエスチョン

  • RQ1最初の投稿に対して返信を受けたユーザーは、憎悪系サブレdditでの継続的参加の可能性が増加するか、減少するか?
  • RQ2返信の感情的・言語的特徴(有毒性、否定的態度、攻撃的発言)は、憎悪コミュニティにおけるユーザーの残留にどのように影響するか?
  • RQ3憎悪系サブレdditにおける返信の影響は、非憎悪コミュニティにおけるそれと比べてどう異なるか?
  • RQ4もし憎悪系サブレdditの最初の返信がより友好的であった場合、参加への累積的影響はどの程度か?
  • RQ5使用された言語モデル(Perspective API、VADER)は、憎悪コミュニティの対立的発言を検出するのに信頼できるか?

主な発見

  • 最初の投稿に対して返信を受けたユーザーは、非憎悪コミュニティとは対照的に、憎悪系サブレdditでの活動継続の可能性が著しく低かった。
  • 憎悪系サブレdditの最初の返信は、憎悪スピーチの使用を制御しても、非憎悪コミュニティのそれよりも著しく有毒で否定的かつ攻撃的であった。
  • 最初の返信の有毒性、コメント投稿者への攻撃、否定的態度の度合いが高いほど、憎悪系サブレdditにおけるユーザー参加の確率が著しく低下する強い負の相関関係が確認された。
  • シミュレーションから、攻撃的返信を中立的または友好的なものに置き換えることで、憎悪系サブレdditにおける長期的参加が著しく増加することが示された。これは、敵対的態度が自己制限的メカニズムとして機能している可能性を示唆している。
  • 非憎悪コミュニティでは、同じシミュレーションでも参加にほとんど変化がなく、これはすでに有毒性や否定的態度が低く抑えられていたためである。
  • Perspective APIの「コメント投稿者への攻撃」および「有毒性」の指標は、それぞれAUCスコア0.85および0.86を達成し、本研究の文脈において信頼性があることが確認された。
Figure 2: Replies to comments in hateful subreddits lead to significantly less engagement than replies to comments in non-hateful subreddits. Distributions of engagement risk ratios for different subreddit types, separated by users who make comments as their first post (A) and users who make submiss
Figure 2: Replies to comments in hateful subreddits lead to significantly less engagement than replies to comments in non-hateful subreddits. Distributions of engagement risk ratios for different subreddit types, separated by users who make comments as their first post (A) and users who make submiss

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。