Skip to main content
QUICK REVIEW

[Paper Review] Developing a Multilingual Annotated Corpus of Misogyny and Aggression

Shiladitya Bhattacharya, Siddharth Singh|arXiv (Cornell University)|Mar 16, 2020
Hate Speech and Cyberbullying DetectionComputer Science34 references61 citations
TL;DR

The paper presents a multilingual annotated corpus of misogyny and aggression in Indian English, Hindi, and Indian Bangla, collected from YouTube comments and annotated along two levels (aggression and misogyny). It covers data collection, annotation tagset, challenges, and baseline classifier experiments across three languages.

ABSTRACT

In this paper, we discuss the development of a multilingual annotated corpus of misogyny and aggression in Indian English, Hindi, and Indian Bangla as part of a project on studying and automatically identifying misogyny and communalism on social media (the ComMA Project). The dataset is collected from comments on YouTube videos and currently contains a total of over 20,000 comments. The comments are annotated at two levels - aggression (overtly aggressive, covertly aggressive, and non-aggressive) and misogyny (gendered and non-gendered). We describe the process of data collection, the tagset used for annotation, and issues and challenges faced during the process of annotation. Finally, we discuss the results of the baseline experiments conducted to develop a classifier for misogyny in the three languages.

Motivation & Objective

  • Motivate the creation of a multilingual corpus to study misogyny and communalism on social media in Indian languages.
  • Describe data collection from YouTube comments across three languages.
  • Define annotation tagsets for aggression (overt, covert, non-aggressive) and misogyny (gendered, non-gendered).
  • Discuss challenges and guidelines in the annotation process.
  • Provide baseline classification results for misogyny detection in the three languages.

Proposed method

  • Collect comments from YouTube videos in Indian English, Hindi, and Indian Bangla.
  • Annotate comments at two levels: aggression (overtly, covertly, non-aggressive) and misogyny (gendered, non-gendered).
  • Describe the tagset and annotation guidelines used byAnnotators.
  • Discuss data collection workflow, quality control, and inter-annotator agreement considerations.
  • Train and report baseline classifiers for misogyny detection across the three languages.

Experimental results

Research questions

  • RQ1How can a multilingual annotated corpus be constructed to study misogyny and aggression in Indian social media?
  • RQ2What are effective annotation schemes for capturing aggression and misogyny across Indian English, Hindi, and Indian Bangla?
  • RQ3What baseline performance can be achieved for misogyny classification in three languages using this corpus?
  • RQ4What challenges arise in collecting and annotating multilingual social media data for this domain?

Key findings

  • The dataset comprises over 20,000 comments annotated for aggression and misogyny across three languages.
  • Aggression is categorized as overt, covert, or non-aggressive; misogyny is categorized as gendered or non-gendered.
  • Baseline experiments were conducted to develop a classifier for misogyny in the three languages.
  • The paper discusses data collection, tagset design, and annotation challenges that impact corpus quality.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.