[论文解读] How is Your Mood When Writing Sexist tweets? Detecting the Emotion Type and Intensity of Emotion Using Natural Language Processing Techniques
本文提出了一种新颖的NLP方法,利用SemEval-2018 Affect in Tweets数据集检测性别歧视推文中的情绪类型与强度。研究识别出四种性别歧视推文类别——间接骚扰、信息威胁、性骚扰和身体骚扰——中的独特情绪特征,表明不同形式的性别歧视与独特的情感状态相关联,这是首次对性别歧视内容中情感状态进行深入分析的研究。
Online social platforms have been the battlefield of users with different emotions and attitudes toward each other in recent years. While sexism has been considered as a category of hateful speech in the literature, there is no comprehensive definition and category of sexism attracting natural language processing techniques. Categorizing sexism as either benevolent or hostile sexism is so broad that it easily ignores the other categories of sexism on social media. Sharifirad S and Matwin S 2018 proposed a well-defined category of sexism including indirect harassment, information threat, sexual harassment and physical harassment, inspired from social science for the purpose of natural language processing techniques. In this article, we take advantage of a newly released dataset in SemEval-2018 task1: Affect in tweets, to show the type of emotion and intensity of emotion in each category. We train, test and evaluate different classification methods on the SemEval- 2018 dataset and choose the classifier with highest accuracy for testing on each category of sexist tweets to know the mental state and the affectual state of the user who tweets in each category. It is a nice avenue to explore because not all the tweets are directly sexist and they carry different emotions from the users. This is the first work experimenting on affect detection this in depth on sexist tweets. Based on our best knowledge they are all new contributions to the field; we are the first to demonstrate the power of such in-depth sentiment analysis on the sexist tweets.
研究动机与目标
- 探究发布性别歧视推文的用户所处的情感状态,超越仇恨言论的二元分类。
- 开发并评估能够检测性别歧视网络内容中情绪类型与强度的NLP模型。
- 探究不同类别的性别歧视(如敌意型与间接型)是否与独特的情绪特征相关。
- 为在线仇恨言论中,特别是性别歧视语境下的情感状态,提供一种新的精细化理解。
- 为情感感知的检测系统奠定基础,使其能更有效地识别并响应社交媒体中的性别歧视内容。
提出的方法
- 利用SemEval-2018 Task 1: Affect in Tweets数据集进行推文中的情感检测。
- 训练并评估多种分类模型,以识别情绪类型(如愤怒、恐惧、喜悦)和强度(低、中、高)。
- 应用多标签分类框架,检测每条推文中的多种情绪,以捕捉情绪的复杂性。
- 将检测到的情绪映射到四种预定义的性别歧视类别:间接骚扰、信息威胁、性骚扰和身体骚扰。
- 基于准确率选择表现最佳的分类器,用于对每种子女歧视推文类别的后续分析。
- 采用统计与定性分析方法,将情绪特征与特定形式的性别歧视表达相关联。
实验结果
研究问题
- RQ1在不同类别的性别歧视推文中,最常表达的情绪类型是什么?
- RQ2不同形式的性别歧视内容中,情绪强度如何变化?
- RQ3是否存在与特定类型性别歧视行为(如间接型与身体骚扰)相关联的独特情绪特征?
- RQ4NLP模型能否有效检测社交媒体性别歧视内容中的情绪类型与强度?
- RQ5发布明显敌意型性别歧视内容的用户与表达更隐蔽形式性别歧视的用户之间,其情感状态差异有多大?
主要发现
- 不同类别的性别歧视推文与独特的情绪特征相关联,表明并非所有性别歧视内容都源于相同的情感状态。
- 在性别歧视推文中,最普遍的情绪是愤怒与恐惧,其强度水平因性别歧视形式而异。
- 被归类为身体骚扰的推文表现出最高强度的负面情绪,尤其是愤怒与恐惧。
- 间接骚扰与信息威胁类别表现出中等到高强度的情绪,表明存在策略性或威胁性意图。
- 表现最佳的分类器在情绪类型与强度检测方面达到高准确率,验证了在性别歧视内容中进行精细化情感检测的可行性。
- 本研究证明,情感感知的NLP模型能够揭示性别歧视网络行为中的细微心理模式,为内容审核与干预提供新见解。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。