Skip to main content
QUICK REVIEW

[논문 리뷰] No Love Among Haters: Negative Interactions Reduce Hate Community Engagement

Daniel Hickey, Matheus Schmitz|arXiv (Cornell University)|2023. 03. 23.
Hate Speech and Cyberbullying Detection인용 수 4
한 줄 요약

이 연구는 인과적 추론을 사용하여 부정적 상호작용—특히 독성, 공격적, 적대적인 댓글—이 악성 레딧 커뮤니티에서 신규 사용자의 참여를 감소시켜, 모더레이션 부족에도 불구하고 모집을 제한한다는 점을 보여준다. 퍼스펙티브 API와 VADER를 사용하여 이러한 적대성이 참가를 저해함을 확인했으며, 시뮬레이션을 통해 친절한 댓글이 악성 서브레딧에서 참여를 크게 증가시킬 수 있음을 확인했다. 반면 비악성 커뮤니티는 안정성을 유지한다.

ABSTRACT

While online hate groups pose significant risks to the health of online platforms and safety of marginalized groups, little is known about what causes users to become active in hate groups and the effect of social interactions on furthering their engagement. We address this gap by first developing tools to find hate communities within Reddit, and then augment 11 subreddits extracted with 14 known hateful subreddits (25 in total). Using causal inference methods, we evaluate the effect of replies on engagement in hateful subreddits by comparing users who receive replies to their first comment (the treatment) to equivalent control users who do not. We find users who receive replies are less likely to become engaged in hateful subreddits than users who do not, while the opposite effect is observed for a matched sample of similar-sized non-hateful subreddits. Using the Google Perspective API and VADER, we discover that hateful community first-repliers are more toxic, negative, and attack the posters more often than non-hateful first-repliers. In addition, we uncover a negative correlation between engagement and attacks or toxicity of first-repliers. We simulate the cumulative engagement of hateful and non-hateful subreddits under the contra-positive scenario of friendly first-replies, finding that attacks dramatically reduce engagement in hateful subreddits. These results counter-intuitively imply that, although under-moderated communities allow hate to fester, the resulting environment is such that direct social interaction does not encourage further participation, thus endogenously constraining the harmful role that these communities could play as recruitment venues for antisocial beliefs.

연구 동기 및 목표

  • 온라인 악성 커뮤니티에서 사용자 참여를 이끄는 요인을 이해하기 위해, 특히 사회적 상호작용의 역할을 파악한다.
  • 초기 상호작용—특히 댓글—이 신규 사용자가 악성 서브레딧에서 지속적으로 참여할 가능성에 어떤 영향을 미치는지 조사한다.
  • 악성 서브레딧과 비악성 커뮤니티에서의 부정적 상호작용의 영향을 비교한다.
  • 댓글의 독성, 공격성, 부정성과 사용자 참여 감소 간의 상관관계를 정량화한다.
  • 친절한 첫 번째 댓글이 악성 커뮤니티에서 장기적 참여에 어떤 영향을 미치는지 시뮬레이션한다.

제안 방법

  • 기존의 악성 서브레딧을 학습 데이터로 사용하여, Reddit 데이터에서 악성 서브레딧을 탐지하는 새로운 방법을 개발하였다.
  • 치료군(첫 번째 게시물에 댓글을 받은 사용자)과 대조군(댓글을 받지 않은 유사 사용자)을 비교하여 인과적 추론을 적용하였다.
  • Google의 퍼스펙티브 API와 VADER를 사용하여 댓글의 독성, 댓글 작성자에 대한 공격성, 부정적 정서를 분석하였다.
  • 첫 번째 댓글이 독성이 없고 공격적이지 않은 '반대 긍정' 시나리오에서 누적 참여도를 추정하기 위한 시뮬레이션 모델을 구축하였다.
  • 유사한 크기와 활동 수준을 가진 서브레딧을 비교하여 사전에 존재하는 혐오 발언 수준을 통제하였다.
  • 인간 레이블링 데이터를 사용하여 퍼스펙티브 API 지표를 검증하였으며, 댓글 공격성에 대해 AUC 0.85, 독성에 대해 AUC 0.86의 성능을 기록하였다.
Figure 1: Schematic of hateful subreddit growth simulation. A mixed effect logistic regression model predicts whether a user continues posting or leaves the subreddit. Predictions are counted to calculate the cumulative number of engaged users in a subreddit.
Figure 1: Schematic of hateful subreddit growth simulation. A mixed effect logistic regression model predicts whether a user continues posting or leaves the subreddit. Predictions are counted to calculate the cumulative number of engaged users in a subreddit.

실험 결과

연구 질문

  • RQ1첫 번째 게시물에 댓글을 받는 것은 악성 서브레딧에서의 지속적 참여 가능성에 증가시키는가, 감소시키는가?
  • RQ2댓글의 정서적 및 언어적 특성(독성, 부정성, 공격성)이 악성 커뮤니티에서 사용자 유지를 어떻게 영향을 미치는가?
  • RQ3악성 서브레딧에서의 댓글 영향은 비악성 서브레딧과 비교해 어떻게 다를까?
  • RQ4악성 서브레딧의 첫 번째 댓글이 더 우호적이면 참여에 어떤 누적 효과가 나타날까?
  • RQ5사용된 언어 모델(Perspective API, VADER)은 악성 커뮤니티 대화에서 적대감을 탐지하는 데 신뢰할 수 있는가?

주요 결과

  • 첫 번째 게시물에 댓글을 받은 사용자는 비악성 커뮤니티의 추세와는 반대로, 악성 서브레딧에서 활동을 지속할 가능성이 뚜렷이 낮았다.
  • 악성 서브레딧의 첫 번째 댓글은 비악성 서브레딧보다 독성이 더 강하고, 더 부정적이며, 더 공격적이었으며, 혐오 발언 사용을 통제한 후에도 여전히 그러했다.
  • 첫 번째 댓글의 독성, 댓글 작성자에 대한 공격성, 부정성과 악성 서브레딧에서의 사용자 참여 확률 사이에 강한 음의 상관관계가 있었다.
  • 시뮬레이션 결과, 적대적인 댓글을 중립적이거나 우호적인 것으로 대체할 경우, 악성 서브레딧에서 장기적 참여가 크게 증가할 것으로 나타났으며, 이는 적대성이 자가 제한 메커니즘으로 작용한다는 것을 시사한다.
  • 비악성 서브레딧은 같은 시뮬레이션 조건에서도 참여에 거의 변화가 없었으며, 이미 낮은 수준의 독성과 부정성 덕분이었다.
  • 퍼스펙티브 API의 댓글 공격성 및 독성 지표는 각각 AUC 0.85와 0.86의 성능을 기록하여, 이 맥락에서의 신뢰성 확인되었다.
Figure 2: Replies to comments in hateful subreddits lead to significantly less engagement than replies to comments in non-hateful subreddits. Distributions of engagement risk ratios for different subreddit types, separated by users who make comments as their first post (A) and users who make submiss
Figure 2: Replies to comments in hateful subreddits lead to significantly less engagement than replies to comments in non-hateful subreddits. Distributions of engagement risk ratios for different subreddit types, separated by users who make comments as their first post (A) and users who make submiss

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.