[论文解读] Developing a Multilingual Annotated Corpus of Misogyny and Aggression
本文提出了一个多语言标注语料库,包含对印度英语、印地语和印度孟加拉语中的厌恶女性言论与攻击性言论的标注,数据来自 YouTube 评论,并在两个层面进行标注(攻击性与厌恶女性言论)。它涵盖数据收集、标注标签集、挑战,以及在三种语言上的基线分类器实验。
In this paper, we discuss the development of a multilingual annotated corpus of misogyny and aggression in Indian English, Hindi, and Indian Bangla as part of a project on studying and automatically identifying misogyny and communalism on social media (the ComMA Project). The dataset is collected from comments on YouTube videos and currently contains a total of over 20,000 comments. The comments are annotated at two levels - aggression (overtly aggressive, covertly aggressive, and non-aggressive) and misogyny (gendered and non-gendered). We describe the process of data collection, the tagset used for annotation, and issues and challenges faced during the process of annotation. Finally, we discuss the results of the baseline experiments conducted to develop a classifier for misogyny in the three languages.
研究动机与目标
- 促使创建一个多语言语料库,以研究印度语言社交媒体上的厌恶女性言论与群体主义。
- 描述来自三种语言的 YouTube 评论的数据收集。
- 为攻击性(明显、隐性、非攻击性)和厌恶女性言论(性别化、非性别化)定义标注标签集。
- 讨论标注过程中的挑战与准则。
- 提供三种语言中厌恶女性言论检测的基线分类结果。
提出的方法
- 从印度英语、印地语和印度孟加拉语的 YouTube 视频中收集评论。
- 在两个层面对评论进行标注:攻击性(显性、隐性、非攻击性)和厌恶女性言论(性别化、非性别化)。
- 描述Annotators使用的标签集和标注准则。
- 讨论数据收集工作流、质量控制和评注者间一致性等考量。
- 训练并报告三种语言的厌恶女性言论检测基线分类器。
实验结果
研究问题
- RQ1如何构建一个多语言标注语料库以研究印度社交媒体中的厌恶女性言论与攻击性?
- RQ2在印度英语、印地语和印度孟加拉语之间,哪些标注方案能够有效捕捉攻击性与厌恶女性言论?
- RQ3使用该语料库在三种语言中进行厌恶女性言论分类可以达到怎样的基线性能?
- RQ4在为该领域收集和标注多语言社交媒体数据时会遇到哪些挑战?
主要发现
- 数据集包含超过 20,000 条评论,跨三种语言对攻击性与厌恶女性言论进行了标注。
- 攻击性被分类为显性、隐性或非攻击性;厌恶女性言论被分类为性别化或非性别化。
- 开展了基线实验以在三种语言中开发厌恶女性言论的分类器。
- 本文讨论了数据收集、标签设计以及影响语料质量的标注挑战。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。