[论文解读] Bangla Text Dataset and Exploratory Analysis for Online Harassment Detection
本文介绍了一个公开可用的孟加拉语文本数据集,包含从公众人物公共主页收集的44,001条Facebook评论,这些评论已针对网络欺凌行为进行标注。研究对多种网络欺凌类别进行了探索性分析,解决了多标签孟加拉语自然语言处理资源稀缺的问题,并推动了在低资源语言如孟加拉语中检测网络欺凌和不当内容的研究。
Being the seventh most spoken language in the world, the use of the Bangla language online has increased in recent times. Hence, it has become very important to analyze Bangla text data to maintain a safe and harassment-free online place. The data that has been made accessible in this article has been gathered and marked from the comments of people in public posts by celebrities, government officials, athletes on Facebook. The total amount of collected comments is 44001. The dataset is compiled with the aim of developing the ability of machines to differentiate whether a comment is a bully expression or not with the help of Natural Language Processing and to what extent it is improper if it is an inappropriate comment. The comments are labeled with different categories of harassment. Exploratory analysis from different perspectives is also included in this paper to have a detailed overview. Due to the scarcity of data collection of categorized Bengali language comments, this dataset can have a significant role for research in detecting bully words, identifying inappropriate comments, detecting different categories of Bengali bullies, etc. The dataset is publicly available at https://data.mendeley.com/datasets/9xjx8twk8p.
研究动机与目标
- 解决孟加拉语网络欺凌检测任务中缺乏标注文本数据集的问题。
- 收集并标注真实世界中孟加拉语的Facebook评论,用于网络欺凌和不当内容的研究。
- 提供一个公开可访问的数据集,以支持孟加拉语等低资源语言的自然语言处理研究。
- 对孟加拉语网络话语中的欺凌行为模式和语言特征进行探索性分析。
- 支持开发能够检测和分类孟加拉语中不同类型网络欺凌行为的机器学习模型。
提出的方法
- 从孟加拉国名人、政府官员和运动员的公共Facebook帖子中抓取评论。
- 该数据集包含44,001条经人工标注的孟加拉语评论,涵盖多个网络欺凌类别。
- 标注采用多类别方案,以区分不同类型的网络欺凌,如仇恨言论、威胁和人身攻击。
- 通过统计摘要、频次分布和可视化(表格与图表)进行探索性数据分析。
- 将数据集公开发布,以支持低资源语言自然语言处理研究的可复现性。
- 分析结合了语言学和社会语言学视角,以理解攻击性语言使用的模式。
实验结果
研究问题
- RQ1在孟加拉语社交媒体评论中,网络欺凌的普遍形式和类别是什么?
- RQ2语言特征和评论特征如何与孟加拉语中不同类型的网络欺凌相关联?
- RQ3公开可用的多标签孟加拉语数据集在多大程度上能支持自动化网络欺凌检测系统的发展?
- RQ4在孟加拉语网络话语中,不同人口统计或内容背景下的攻击性语言分布模式如何?
- RQ5对低资源语言数据集的探索性分析如何为未来孟加拉语自然语言处理模型设计提供指导?
主要发现
- 该数据集包含44,001条孟加拉语评论,涵盖多个网络欺凌类别,为自然语言处理研究提供了丰富资源。
- 探索性分析揭示了不同网络欺凌类型中攻击性表达的显著语言模式和频次分布特征。
- 该数据集显示出网络欺凌在严重程度和性质上的显著差异,包括仇恨言论、人身攻击和威胁。
- 该数据集的可用性填补了孟加拉语自然语言处理领域的一项关键空白,尤其在在线安全和内容审核应用方面。
- 本研究凸显了在大规模真实世界场景下对孟加拉语文本进行标注的可行性,为未来模型的训练与评估提供了支持。
- 该数据集可通过提供的URL公开获取,支持研究的可复现性与社区驱动的研究发展。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。