[Paper Review] Antisemitic Messages? A Guide to High-Quality Annotation and a Labeled Dataset of Tweets
This paper presents a high-quality annotation framework and a labeled dataset of 6,941 English tweets (18% antisemitic) from January 2019 to December 2021, using the IHRA definition of antisemitism. The method enforces strict application of the definition by requiring annotators to specify which part applies and allows personal disagreement, reducing false positives in automated detection.
One of the major challenges in automatic hate speech detection is the lack of datasets that cover a wide range of biased and unbiased messages and that are consistently labeled. We propose a labeling procedure that addresses some of the common weaknesses of labeled datasets. We focus on antisemitic speech on Twitter and create a labeled dataset of 6,941 tweets that cover a wide range of topics common in conversations about Jews, Israel, and antisemitism between January 2019 and December 2021 by drawing from representative samples with relevant keywords. Our annotation process aims to strictly apply a commonly used definition of antisemitism by forcing annotators to specify which part of the definition applies, and by giving them the option to personally disagree with the definition on a case-by-case basis. Labeling tweets that call out antisemitism, report antisemitism, or are otherwise related to antisemitism (such as the Holocaust) but are not actually antisemitic can help reduce false positives in automated detection. The dataset includes 1,250 tweets (18%) that are antisemitic according to the International Holocaust Remembrance Alliance (IHRA) definition of antisemitism. It is important to note, however, that the dataset is not comprehensive. Many topics are still not covered, and it only includes tweets collected from Twitter between January 2019 and December 2021. Additionally, the dataset only includes tweets that were written in English. Despite these limitations, we hope that this is a meaningful contribution to improving the automated detection of antisemitic speech.
Motivation & Objective
- To address the lack of consistently labeled datasets covering a broad range of biased and unbiased messages related to antisemitism.
- To reduce false positives in automated antisemitic speech detection by including non-antisemitic tweets that discuss antisemitism, the Holocaust, or Israel.
- To develop a rigorous annotation process that enforces adherence to the IHRA definition of antisemitism through structured labeling.
- To create a representative, topic-diverse dataset of tweets from a specific time window (2019–2021) in English.
- To contribute a methodological guide for high-quality annotation that supports reproducibility and consistency in hate speech research.
Proposed method
- Annotators were required to specify which component of the IHRA definition of antisemitism applied to each tweet.
- Annotators had the option to personally disagree with the definition’s application on a case-by-case basis, ensuring transparency and reducing bias.
- Tweets were collected using relevant keywords from a representative sample between January 2019 and December 2021.
- The dataset includes 1,250 antisemitic tweets (18%) according to the IHRA definition, with the remainder labeled as non-antisemitic.
- The annotation process emphasized consistency and clarity by structuring labeling around a standardized definition.
- The dataset was curated to include tweets that report, critique, or discuss antisemitism without being antisemitic themselves.
Experimental results
Research questions
- RQ1How can annotation processes be improved to ensure consistent and reliable labeling of antisemitic content?
- RQ2What proportion of tweets discussing Jews, Israel, or the Holocaust are actually antisemitic according to the IHRA definition?
- RQ3Can including non-antisemitic but antisemitism-related content improve the performance of automated antisemitic speech detection systems?
- RQ4To what extent does a structured annotation process that ties labels to specific definition components reduce labeling ambiguity?
- RQ5How representative is a dataset of English tweets from 2019–2021 in capturing the full spectrum of antisemitic and non-antisemitic discourse?
Key findings
- The dataset contains 6,941 tweets collected from Twitter between January 2019 and December 2021, all in English.
- Of these, 1,250 tweets (18%) were classified as antisemitic according to the IHRA definition.
- The annotation process successfully applied the IHRA definition by requiring annotators to identify which specific component of the definition applied to each tweet.
- The inclusion of tweets that discuss or report antisemitism but are not antisemitic themselves helps reduce false positives in automated detection systems.
- The dataset is not comprehensive, as it excludes non-English content and covers only a subset of topics related to antisemitism.
- Despite limitations, the dataset and annotation framework represent a meaningful contribution to improving the accuracy and consistency of antisemitic speech detection.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.