Skip to main content
QUICK REVIEW

[论文解读] Annotating Antisemitic Online Content. Towards an Applicable Definition of Antisemitism

Günther Jikeli, Damir Ćavar|arXiv (Cornell University)|Sep 29, 2019
Hate Speech and Cyberbullying Detection参考文献 31被引用 8
一句话总结

本文提出了一种上下文敏感的标注框架,用于使用国际大屠杀纪念联盟(IHRA)的反犹主义定义,在在线平台中识别反犹主义内容,研究对象为随机抽取的推特数据。研究结果表明,在关于犹太人和以色列的推特对话中,超过10%的内容属于反犹主义或可能属于反犹主义,因此需要接受过时事培训的专家标注员,以确保检测的准确性和一致性。

ABSTRACT

Online antisemitism is hard to quantify. How can it be measured in rapidly growing and diversifying platforms? Are the numbers of antisemitic messages rising proportionally to other content or is it the case that the share of antisemitic content is increasing? How does such content travel and what are reactions to it? How widespread is online Jew-hatred beyond infamous websites and fora, and closed social media groups? However, at the root of many methodological questions is the challenge of finding a consistent way to identify diverse manifestations of antisemitism in large datasets. What is more, a clear definition is essential for building an annotated corpus that can be used as a gold standard for machine learning programs to detect antisemitic online content. We argue that antisemitic content has distinct features that are not captured adequately in generic approaches of annotation, such as hate speech, abusive language, or toxic language. We discuss our experiences with annotating samples from our dataset that draw on a ten percent random sample of public tweets from Twitter. We show that the widely used definition of antisemitism by the International Holocaust Remembrance Alliance can be applied successfully to online messages if inferences are spelled out in detail and if the focus is not on intent of the disseminator but on the message in its context. However, annotators have to be highly trained and knowledgeable about current events to understand each tweet's underlying message within its context. The tentative results of the annotation of two of our small but randomly chosen samples suggest that more than ten percent of conversations on Twitter about Jews and Israel are antisemitic or probably antisemitic. They also show that at least in conversations about Jews, an equally high number of tweets denounce antisemitism, although these conversations do not necessarily coincide.

研究动机与目标

  • 开发一种可靠的方法,用于识别大规模在线数据集中的反犹主义内容。
  • 评估IHRA反犹主义定义在多样化在线信息中的适用性,特别是在社交媒体语境中。
  • 创建一个高质量标注语料库,用于训练机器学习模型以检测反犹主义。
  • 研究公共推特对话中关于犹太人和以色列的反犹主义内容的普遍程度及其传播情况。
  • 解决通用仇恨言论或有毒语言标注框架在捕捉反犹主义细微差别的局限性。

提出的方法

  • 本研究将国际大屠杀纪念联盟(IHRA)的反犹主义定义应用于关于犹太人和以色列的公共推特推文的10%随机样本。
  • 标注员接受过训练,基于上下文线索解读信息,重点关注内容的影响而非发布者的意图。
  • 标注过程需要深入理解上下文,包括对时事和历史引用的知识。
  • 采用两阶段标注流程:先进行初步标注,随后由专家审查以解决模糊性问题。
  • 对数据集进行反犹主义内容分析,重点识别语言、象征符号和话语模式。
  • 通过标注者间一致性检查和标注指南的迭代优化对框架进行验证。

实验结果

研究问题

  • RQ1IHRA反犹主义定义能否在在线内容中一致应用,特别是在动态且模糊的社交媒体环境中?
  • RQ2在关于犹太人和以色列的公共推特对话中,有多少比例包含反犹主义内容?
  • RQ3反犹主义信息在语言和语境特征上与一般仇恨言论或侮辱性语言有何不同?
  • RQ4在这些对话中,用户在多大程度上谴责反犹主义?此类谴责与反犹主义内容如何共现?
  • RQ5需要何种程度的标注员专业知识才能可靠地识别上下文中的反犹主义内容?

主要发现

  • 根据IHRA定义,超过10%的关于犹太人和以色列的推特对话被归类为反犹主义或可能属于反犹主义。
  • 同一数据集显示,这些对话中同样比例的推文积极谴责反犹主义,表明话语存在高度极化。
  • 研究发现,准确标注依赖于对上下文的深入理解,因为许多反犹主义信息依赖于隐晦语言、历史引用或反讽。
  • 为实现可靠的标注者间一致性,标注员需接受时事和反犹主义刻板印象的专门培训。
  • 当结合详细的上下文推断时,IHRA定义在在线内容中具有适用性,但需要仔细解读。
  • 结果表明,反犹主义在公共在线话语中的普遍存在程度高于以往量化结果,尤其是在涉及以色列的讨论中。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。