Skip to main content
QUICK REVIEW

[Paper Review] Dataset for Identification of Homophobia and Transophobia in Multilingual YouTube Comments

Bharathi Raja Chakravarthi, Ruba Priyadharshini|arXiv (Cornell University)|Sep 1, 2021
Hate Speech and Cyberbullying Detection64 references56 citations
TL;DR

This paper presents a hierarchical taxonomy and an expert-labeled multilingual YouTube comments dataset to identify homophobia and transphobia, along with baseline models.

ABSTRACT

The increased proliferation of abusive content on social media platforms has a negative impact on online users. The dread, dislike, discomfort, or mistrust of lesbian, gay, transgender or bisexual persons is defined as homophobia/transphobia. Homophobic/transphobic speech is a type of offensive language that may be summarized as hate speech directed toward LGBT+ people, and it has been a growing concern in recent years. Online homophobia/transphobia is a severe societal problem that can make online platforms poisonous and unwelcome to LGBT+ people while also attempting to eliminate equality, diversity, and inclusion. We provide a new hierarchical taxonomy for online homophobia and transphobia, as well as an expert-labelled dataset that will allow homophobic/transphobic content to be automatically identified. We educated annotators and supplied them with comprehensive annotation rules because this is a sensitive issue, and we previously discovered that untrained crowdsourcing annotators struggle with diagnosing homophobia due to cultural and other prejudices. The dataset comprises 15,141 annotated multilingual comments. This paper describes the process of building the dataset, qualitative analysis of data, and inter-annotator agreement. In addition, we create baseline models for the dataset. To the best of our knowledge, our dataset is the first such dataset created. Warning: This paper contains explicit statements of homophobia, transphobia, stereotypes which may be distressing to some readers.

Motivation & Objective

  • Propose a hierarchical taxonomy for online homophobia and transphobia.
  • Create and share an expert-labeled multilingual dataset of YouTube comments.
  • Ensure annotation quality through educator-led guidelines due to cultural sensitivities.
  • Evaluate inter-annotator agreement on the labeling process.
  • Provide baseline models to identify homophobic/transphobic content.

Proposed method

  • Develop a new hierarchical taxonomy for homophobia and transphobia in online comments.
  • Assemble and annotate a multilingual dataset with expert annotators and comprehensive rules.
  • Educate annotators and use structured annotation guidelines to mitigate biases.
  • Analyze qualitative aspects and inter-annotator agreement of annotations.
  • Construct baseline models for automatic identification of targeted content.

Experimental results

Research questions

  • RQ1What constitutes homophobia and transphobia in multilingual online YouTube comments under a hierarchical taxonomy?
  • RQ2How can expert annotation and clear rules improve labeling reliability for sensitive content?
  • RQ3What is the size and multilingual composition of the annotated dataset?
  • RQ4How do baseline models perform on identifying homophobic/transphobic content in multilingual YouTube comments?
  • RQ5What is the inter-annotator agreement in the annotation process?

Key findings

  • The dataset contains 15,141 annotated multilingual comments.
  • An expert-led annotation process with comprehensive rules was used to improve reliability.
  • The paper analyzes qualitative aspects of the data and reports inter-annotator agreement.
  • Baseline models were created to establish initial performance on the dataset.
  • This work appears to be among the first to provide such a dataset for this topic.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.