Skip to main content
QUICK REVIEW

[논문 리뷰] Annotating Antisemitic Online Content. Towards an Applicable Definition of Antisemitism

Günther Jikeli, Damir Ćavar|arXiv (Cornell University)|2019. 09. 29.
Hate Speech and Cyberbullying Detection참고 문헌 31인용 수 8
한 줄 요약

이 논문은 온라인 플랫폼에서 유대인에 대한 반유대주의 콘텐츠를 식별하기 위해 IHRA의 반유대주의 정의를 적용한 맥락 민감성 주석 프레임워크를 제안한다. 무작위로 선정된 트위터 데이터에 적용한 결과, 유대인과 이스라엘에 관한 대화 중 10퍼센트 이상이 반유대주의이거나 반유대주의로 간주될 가능성이 있음을 입증하였다. 이는 정확성과 일관성을 확보하기 위해 최신 사건에 익숙한 전문 주석자가 필요하다는 것을 시사한다.

ABSTRACT

Online antisemitism is hard to quantify. How can it be measured in rapidly growing and diversifying platforms? Are the numbers of antisemitic messages rising proportionally to other content or is it the case that the share of antisemitic content is increasing? How does such content travel and what are reactions to it? How widespread is online Jew-hatred beyond infamous websites and fora, and closed social media groups? However, at the root of many methodological questions is the challenge of finding a consistent way to identify diverse manifestations of antisemitism in large datasets. What is more, a clear definition is essential for building an annotated corpus that can be used as a gold standard for machine learning programs to detect antisemitic online content. We argue that antisemitic content has distinct features that are not captured adequately in generic approaches of annotation, such as hate speech, abusive language, or toxic language. We discuss our experiences with annotating samples from our dataset that draw on a ten percent random sample of public tweets from Twitter. We show that the widely used definition of antisemitism by the International Holocaust Remembrance Alliance can be applied successfully to online messages if inferences are spelled out in detail and if the focus is not on intent of the disseminator but on the message in its context. However, annotators have to be highly trained and knowledgeable about current events to understand each tweet's underlying message within its context. The tentative results of the annotation of two of our small but randomly chosen samples suggest that more than ten percent of conversations on Twitter about Jews and Israel are antisemitic or probably antisemitic. They also show that at least in conversations about Jews, an equally high number of tweets denounce antisemitism, although these conversations do not necessarily coincide.

연구 동기 및 목표

  • 대규모 온라인 데이터셋에서 반유대주의 콘텐츠를 신뢰성 있게 식별하는 방법을 개발하기 위해.
  • 반유대주의의 IHRA 정의가 다양한 온라인 메시지, 특히 소셜 미디어 환경에서 적용 가능한지 평가하기 위해.
  • 기계 학습 모델이 반유대주의를 탐지하도록 훈련하기 위한 골드 표준 주석 코퍼스를 구축하기 위해.
  • 유대인과 이스라엘에 관한 공개 트위터 대화에서 반유대주의 콘텐츠의 보편성과 확산 양상을 분석하기 위해.
  • 일반적인 혐오 발언 또는 유해 언어 주석 프레임워크의 한계를 극복하여 반유대주의의 미묘한 특징을 포착할 수 있는 방법을 모색하기 위해.

제안 방법

  • 연구는 유대인과 이스라엘에 관해 논의하는 공개 트위터 트윗의 10퍼센트 무작위 샘플에 국제 홀로코스트 기억 협의회(IHRA)의 반유대주의 정의를 적용하였다.
  • 주석자는 메시지의 맥락적 단서를 기반으로 해석하도록 훈련되었으며, 작성자의 의도보다는 메시지의 영향에 초점을 맞추었다.
  • 주석 과정에는 현재 일어나는 사건과 역사적 참조에 대한 지식을 포함한 깊이 있는 맥락 이해가 필요했다.
  • 이중 단계 주석 절차를 사용하였는데, 초기 레이블링 이후 전문가 검토를 통해 모호성을 해소하였다.
  • 데이터셋은 반유대주의 콘텐츠를 분석하였으며, 언어, 상징, 논의 양상에서의 패턴을 집중적으로 분석하였다.
  • 주석 지침의 반복적 개선과 상호 주석자 간 일致도 검증을 통해 프레임워크를 검증하였다.

실험 결과

연구 질문

  • RQ1IHRA의 반유대주의 정의는 동적인 고도로 모호한 소셜 미디어 환경에서 온라인 콘텐츠에 일관되게 적용될 수 있는가?
  • RQ2유대인과 이스라엘에 관해 논의하는 공개 트위터 대화 중 반유대주의 콘텐츠를 포함하는 비율은 얼마인가?
  • RQ3반유대주의 메시지는 일반적인 혐오 발언이나 폭력적 언어와 비교해 언어적 및 맥락적 특징에서 어떻게 다를까?
  • RQ4이러한 대화에서 반유대주의를 비난하는 사용자는 얼마나 되며, 이러한 비난은 반유대주의 콘텐츠와 동시에 어떻게 발생하는가?
  • RQ5반유대주의 콘텐츠를 맥락에서 신뢰성 있게 식별하기 위해 주석자에게 요구되는 전문성 수준은 어느 정도인가?

주요 결과

  • IHRA 정의에 따라 분석한 결과, 유대인과 이스라엘에 관해 논의하는 트위터 대화 중 10퍼센트 이상이 반유대주의 또는 반유대주의로 간주될 가능성이 있음이 확인되었다.
  • 동일한 데이터셋을 분석한 결과, 이와 유사한 비율의 트윗이 반유대주의를 적극적으로 비난하고 있었으며, 이는 논의의 극심한 분열 수준을 시사한다.
  • 연구는 정확한 주석을 위해서는 맥락 이해가 필수적임을 발견하였으며, 많은 반유대주의 메시지가 암시어, 역사적 참조, 풍자 등을 활용하기 때문이다.
  • 높은 상호 주석자 간 일致도를 달성하기 위해 주석자들은 최신 사건과 반유대주의의 전형적 패턴에 대한 특수 훈련이 필요했다.
  • 세부적인 맥락 인식을 수반할 경우 IHRA 정의는 온라인 콘텐츠에 적용 가능했지만, 신중한 해석이 요구됨을 확인하였다.
  • 결과적으로 반유대주의가 이전에 측정된 것보다 공개 온라인 논의에서 더 광범위하게 퍼져 있다는 점을 시사하며, 특히 이스라엘을 둘러싼 논의에서 더욱 그렇다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.