Skip to main content
QUICK REVIEW

[论文解读] Analyzing the hate and counter speech accounts on Twitter

Binny Mathew, N. Ravi Kumar|arXiv (Cornell University)|Dec 6, 2018
Hate Speech and Cyberbullying Detection参考文献 32被引用 58
一句话总结

该论文构建了一个关于 Twitter 上仇恨推文及其对抗性言论的数据集,分析语言与心理语言学特征,轮廓化仇恨账户与对抗账户,并训练出一个能够区分仇恨账户与对抗账户的分类器,其 F1 为 0.78(78% 的准确率)。

ABSTRACT

The online hate speech is proliferating with several organization and countries implementing laws to ban such harmful speech. While these restrictions might reduce the amount of such hateful content, it does so by restricting freedom of speech. Thus, an promising alternative supported by several organizations is to counter such hate speech with more speech. In this paper, We analyze hate speech and the corresponding counters (aka counterspeech) on Twitter. We perform several lexical, linguistic and psycholinguistic analysis on these user accounts and obverse that counter speakers employ several strategies depending on the target community. The hateful accounts express more negative sentiments and are more profane. We also find that the hate tweets by verified accounts have much more virality as compared to a tweet by a non-verified account. While the hate users seem to use words more about envy, hate, negative emotion, swearing terms, ugliness, the counter users use more words related to government, law, leader. We also build a supervised model for classifying the hateful and counterspeech accounts on Twitter and obtain an F-score of 0.77. We also make our dataset public to help advance the research on hate speech.

研究动机与目标

  • 推动并研究对抗性言论作为在 Twitter 上屏蔽仇恨言论的替代方案。
  • 创建一个仇恨推文及其对抗性回复的数据集用于分析与建模。
  • 在活动性、词汇、个性与话题维度上描述仇恨与对抗性言论账户。
  • 开发一个预测模型以自动区分仇恨账户与对抗性言论账户。
  • 提供关于对抗性言论策略如何因目标社区与平台动态而异的见解。

提出的方法

  • 整理一个包含 558 条仇恨推文的 1290 条对抗性回复的数据集,产生来自 1239 个账户的 1290 条对抗性回复以及来自 548 个仇恨账户的 558 条仇恨推文。
  • 对推文进行注释以识别仇恨内容并将对抗性言论分类到预定义类别(并给出评注者间一致性度量)。
  • 提取并分析词汇、情感、粗话与心理语言特征(如 Empath 类别、IBM Watson 人格特质)。
  • 基于 3200 条推文历史构建用户层面的特征(TF-IDF、个人资料指标、词汇/情感特征),用于每个账户。
  • 训练并评估多种分类器(SVM、LR、RF、ET、XGBoost、CatBoost)以区分仇恨与对抗性言论账户,最终选择 CatBoost 作为最佳。
  • 进行特征消融以评估 TF-IDF、词汇和情感特征的贡献。

实验结果

研究问题

  • RQ1仇恨与对抗性言论账户在 Twitter 上的词汇、情感与心理语言学差异是什么?
  • RQ2对抗性言论策略在不同目标社区(如宗教、国籍、种族、性取向)之间有何变化?
  • RQ3账户层面的特征是否能可靠地区分仇恨与对抗性言论账户,哪些特征对预测贡献最大?
  • RQ4区分仇恨与对抗性言论账户的主题兴趣和人格特质有哪些差异?

主要发现

  • 观察到 558 条仇恨推文对应 1290 条对抗性回复,对抗性回复占比为 75.39%。
  • 最佳分类器(CatBoost)在区分仇恨与对抗性言论账户方面达到 78% 的准确率(F1 0.77;XGBoost 达到 74%)。
  • 仇恨账户往往更老、更活跃且拥有更多粉丝,而对抗性言论账户每天拥有更多的朋友。
  • 来自经过验证账户的仇恨推文显示出比未验证仇恨推文更高的传播性;经过验证的仇恨账户获得的参与度指标显著更高。
  • 词汇分析显示仇恨账户使用更多的嫉妒、仇恨、负面情绪与粗话,而对抗性言论账户使用与政府、法律和领导力相关的词汇更多;Empath 和人格分析显示对抗者更高的宜人性,仇恨账户则更高的外向性。
  • 话题分析表明,对抗性言论更关注政治、新闻与新闻业,而仇恨账户则集中在种族侮辱等主题。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。