[论文解读] Anubhuti -- An annotated dataset for emotional analysis of Bengali short stories
本文介绍了 Anubhuti,这是首个针对孟加拉语短篇小说情感分析的大规模标注数据集,通过语言学专业标注员进行严格的人工标注,实现了较高的标注者间一致性。该数据集使传统机器学习和深度学习模型能够实现高精度的情感分类,为低资源语言的自然语言处理、语言学及文学情感分析提供了宝贵见解。
Thousands of short stories and articles are being written in many different languages all around the world today. Bengali, or Bangla, is the second highest spoken language in India after Hindi and is the national language of the country of Bangladesh. This work reports in detail the creation of Anubhuti -- the first and largest text corpus for analyzing emotions expressed by writers of Bengali short stories. We explain the data collection methods, the manual annotation process and the resulting high inter-annotator agreement of the dataset due to the linguistic expertise of the annotators and the clear methodology of labelling followed. We also address some of the challenges faced in the collection of raw data and annotation process of a low resource language like Bengali. We have verified the performance of our dataset with baseline Machine Learning as well as a Deep Learning model for emotion classification and have found that these standard models have a high accuracy and relevant feature selection on Anubhuti. In addition, we also explain how this dataset can be of interest to linguists and data analysts to study the flow of emotions as expressed by writers of Bengali literature.
研究动机与目标
- 创建首个大规模、高质量的孟加拉语短篇小说情感分析标注数据集。
- 解决在孟加拉语等低资源语言中收集和标注文本数据的挑战。
- 通过专家标注员和明确的标注方法,确保高标注者间一致性。
- 利用标准机器学习和深度学习模型评估该数据集在情感分类中的实用性。
- 支持自然语言处理、计算语言学和文学情感流分析等跨学科研究。
提出的方法
- 数据收集通过从多样化的文学来源获取孟加拉语短篇小说,以确保语言和主题的多样性。
- 由接受过语言学训练的标注员使用标准化的标注框架,对文本中表达的情感进行人工标注。
- 通过测量标注者间一致性,发现其水平较高,表明情感标注具有一致性和可靠性。
- 在 Anubhuti 数据集上训练并评估了基线模型,包括传统机器学习和深度学习架构。
- 进行了特征选择,以识别对情感分类有贡献的语言线索。
- 通过情感分类任务的性能评估对数据集进行了验证,结果表明模型具有出色的准确率。
实验结果
研究问题
- RQ1如何为孟加拉语等低资源语言构建一个高质量、大规模的情感标注数据集?
- RQ2在孟加拉语中,使用专家标注员和标准化标注协议可实现多高的标注者间一致性?
- RQ3标准机器学习和深度学习模型在 Anubhuti 数据集上的情感分类表现如何?
- RQ4在孟加拉语文学文本中,哪些语言特征最能预测情感?
- RQ5该数据集在支持自然语言处理、语言学和文学分析等跨学科研究方面有哪些作用?
主要发现
- Anubhuti 数据集实现了高标注者间一致性,证实了标注过程的可靠性。
- 标准机器学习和深度学习模型在 Anubhuti 数据集上进行情感分类时表现出高准确率。
- 该数据集支持有效的特征选择,识别出孟加拉语叙事中关键的情感检测语言线索。
- 该语料库是目前唯一且最大的孟加拉语短篇小说情感分析资源,填补了低资源自然语言处理领域的关键空白。
- 该数据集不仅对自然语言处理任务有价值,也对分析孟加拉语文学中情感流动的语言学和文学研究具有重要意义。
- 该数据集的创建过程为其他低资源语言构建类似资源提供了可复现的框架。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。