Skip to main content
QUICK REVIEW

[论文解读] Counter Hate Speech in Social Media: A Survey

Dana Alsagheer, Hadi Mansourifar|arXiv (Cornell University)|Feb 21, 2022
Hate Speech and Cyberbullying Detection被引用 5
一句话总结

本综述研究了社交媒体中的反仇恨言论(CHS)生成,作为一种主动应对仇恨言论的策略,同时保护言论自由。通过将现有研究按方法论、数据集和影响评估进行分类,揭示了诸如缺乏对评论序列的纵向分析以及CHS干预在现实世界中效果不确定等研究空白。

ABSTRACT

With the high prevalence of offensive language against minorities in social media, counter-hate speeches (CHS) generation is considered an automatic way of tackling this challenge. The CHS is supposed to appear as a third voice to educate people and keep the social [red lines bold] without limiting the principles of freedom of speech. In this paper, we review the most important research in the past and present with a main focus on methodologies, collected datasets and statistical analysis CHS's impact on social media. The CHS generation is based on the optimistic assumption that any attempt to intervene the hate speech in social media can play a positive role in this context. Beyond that, previous works ignored the investigation of the sequence of comments before and after the CHS. However, the positive impact is not guaranteed, as shown in some previous works. To the best of our knowledge, no attempt has been made to survey the related work to compare the past research in terms of CHS's impact on social media. We take the first step in this direction by providing a comprehensive review on related works and categorizing them based on different factors including impact, methodology, data source, etc.

研究动机与目标

  • 应对社交媒体平台上针对少数群体的仇恨言论日益普遍的问题。
  • 探讨自动化反仇恨言论作为第三方干预手段的作用,以教育用户而不压制言论自由。
  • 识别在评估CHS长期影响方面存在的研究空白,特别是干预前后评论序列动态的变化。
  • 对现有CHS生成方法进行全面、系统的综述,涵盖方法论、数据来源和评估指标。
  • 通过比较以往研究在影响、模型设计和伦理考量方面的异同,为未来研究奠定基础。

提出的方法

  • 对2018年至2022年间关于反仇恨言论的同行评审论文和预印本研究进行系统性文献综述。
  • 根据方法论(如序列到序列模型、微调的Transformer架构)、数据来源(如Twitter、Reddit)和评估策略,对CHS研究进行分类。
  • 分析先前研究中报告的统计性能指标,包括仇恨言论检测和CHS生成的F1值、精确率和召回率。
  • 通过分析评论线程动态,评估CHS对社交媒体话语的影响,包括干预前后的感情倾向和用户参与度模式。
  • 识别与先前研究的文本重叠,以确保方法论的一致性,并避免在综述文献中出现重复。
  • 使用结构化分类体系,按任务类型(如响应生成、缓解、教育)和部署情境对CHS系统进行分类。

实验结果

研究问题

  • RQ1反仇恨言论生成中占主导地位的方法论是什么?它们在设计和性能上如何不同?
  • RQ2现有CHS数据集在来源、标注质量和真实社交媒体话语代表性方面有何差异?
  • RQ3CHS对社交媒体线程中后续用户互动和话语质量的可衡量影响是什么?
  • RQ4为何一些CHS干预未能产生积极效果?哪些因素导致其无效或产生意外后果?
  • RQ5在评估自动化CHS对在线社区的长期社会和心理影响方面,关键的研究空白是什么?

主要发现

  • 大多数CHS生成系统依赖于微调的Transformer模型,如BART和T5,在基准数据集上实现了中等F1值(通常高于0.70)。
  • 尽管在孤立指标上表现良好,但CHS在现实世界中的实际影响仍不明确,原因在于对评论线程演变和用户行为变化的评估不足。
  • 许多研究忽略了CHS部署前后评论序列的动态,限制了对话语发展进程及潜在反效果的理解。
  • 本综述与先前研究存在显著的文本重叠(如arXiv:1909.04251、arXiv:2009.08392),表明未来研究需要更具区分度的贡献。
  • 评估协议尚未达成共识,不同研究使用各异的指标,且缺乏用于评估CHS有效性的标准化基准,超越二元仇恨言论检测。
  • 乐观假设CHS始终能改善话语质量,并未得到实证证据的持续支持,部分研究报告其对用户参与度和情感倾向的影响微乎其微,甚至为负面。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。