[論文レビュー] Annotating Antisemitic Online Content. Towards an Applicable Definition of Antisemitism
本稿は、IHRAの反ユダヤ主義の定義を用いて、オンラインプラットフォームにおける反ユダヤ的コンテンツを同定する文脈に敏感なアノテーションフレームワークを提案する。Twitterデータのランダムサンプルに適用した結果、ユダヤ人やイスラエルに関する会話の10%以上が反ユダヤ的またはおそらく反ユダヤ的であることが示された。これは、正確性と一貫性を確保するため、最新の出来事に精通した専門的アノテーターの訓練が不可欠であることを示している。
Online antisemitism is hard to quantify. How can it be measured in rapidly growing and diversifying platforms? Are the numbers of antisemitic messages rising proportionally to other content or is it the case that the share of antisemitic content is increasing? How does such content travel and what are reactions to it? How widespread is online Jew-hatred beyond infamous websites and fora, and closed social media groups? However, at the root of many methodological questions is the challenge of finding a consistent way to identify diverse manifestations of antisemitism in large datasets. What is more, a clear definition is essential for building an annotated corpus that can be used as a gold standard for machine learning programs to detect antisemitic online content. We argue that antisemitic content has distinct features that are not captured adequately in generic approaches of annotation, such as hate speech, abusive language, or toxic language. We discuss our experiences with annotating samples from our dataset that draw on a ten percent random sample of public tweets from Twitter. We show that the widely used definition of antisemitism by the International Holocaust Remembrance Alliance can be applied successfully to online messages if inferences are spelled out in detail and if the focus is not on intent of the disseminator but on the message in its context. However, annotators have to be highly trained and knowledgeable about current events to understand each tweet's underlying message within its context. The tentative results of the annotation of two of our small but randomly chosen samples suggest that more than ten percent of conversations on Twitter about Jews and Israel are antisemitic or probably antisemitic. They also show that at least in conversations about Jews, an equally high number of tweets denounce antisemitism, although these conversations do not necessarily coincide.
研究の動機と目的
- 大規模なオンラインデータセットにおける反ユダヤ的コンテンツを特定する信頼性の高い手法を開発すること。
- 反ユダヤ主義のIHRA定義が、特にソーシャルメディア環境における多様なオンラインメッセージに適用可能かどうかを評価すること。
- 機械学習モデルによる反ユダヤ主義検出のためのゴールドスタンダードアノテートコーパスを構築すること。
- ユダヤ人やイスラエルに関する公共のTwitter会話において、反ユダヤ的コンテンツの広がりと普及状況を調査すること。
- 一般的な嫌がらせや有毒な言語のアノテーションフレームワークの限界を是正し、反ユダヤ主義の微細な特徴を捉えること。
提案手法
- 本研究は、ユダヤ人やイスラエルをテーマとする公共のTwitterツイートの10%ランダムサンプルに、国際ホロコースト記憶連合(IHRA)の反ユダヤ主義の定義を適用した。
- アノテーターは、意図ではなく、メッセージの影響に注目して文脈的ヒントに基づいて解釈するよう訓練された。
- アノテーションプロセスには、現在の出来事や歴史的参照を含む、深い文脈的理解が不可欠であった。
- 二段階のアノテーションプロセスを採用した:初期ラベル付けの後、曖昧さを解消するための専門家レビューを実施した。
- データセットは、言語、象徴、ディス course のパターンを特定することに注力して、反ユダヤ的コンテンツについて分析された。
- 相互アノテーター整合性のチェックとアノテーションガイドラインの反復的改善を通じて、フレームワークの妥当性を検証した。
実験結果
リサーチクエスチョン
- RQ1反ユダヤ主義のIHRA定義は、動的で曖昧なソーシャルメディア環境において、オンラインコンテンツに一貫して適用可能だろうか?
- RQ2ユダヤ人やイスラエルに関する公共のTwitter会話のうち、何パーセントが反ユダヤ的コンテンツを含むだろうか?
- RQ3反ユダヤ的メッセージは、一般の嫌がらせや攻撃的言語と比べて、言語的および文脈的特徴でどのように異なるだろうか?
- RQ4これらの会話において、ユーザーはどれほど反ユダヤ主義を非難しており、その非難は反ユダヤ的コンテンツとどのように共起するだろうか?
- RQ5文脈における反ユダヤ的コンテンツを信頼性高く同定するには、どの程度のアノテーターの専門的知識が求められるだろうか?
主な発見
- IHRA定義に基づく分析により、ユダヤ人やイスラエルに関するTwitter会話の10%以上が反ユダヤ的またはおそらく反ユダヤ的と分類された。
- 同じデータセットにおいて、これらの会話内のツイートのほぼ同程度の割合が反ユダヤ主義を明確に非難しており、議論の両極化が顕著であることが示された。
- 本研究では、多くの反ユダヤ的メッセージが暗号的表現、歴史的参照、あるいは皮肉を用いるため、文脈的理解が正確なアノテーションに不可欠であることが判明した。
- 相互アノテーター整合性を達成するためには、最新の出来事と反ユダヤ的レトリックに精通した特別な訓練が必要であった。
- 詳細な文脈的推論を伴えば、IHRA定義はオンラインコンテンツに適用可能であることが示されたが、解釈には注意が要する。
- 結果から、特にイスラエルを巡る議論において、反ユダヤ主義がこれまでに測定されたよりも広く公的なオンラインディス course に存在している可能性が示唆された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。