[论文解读] Semantic Enrichment of Nigerian Pidgin English for Contextual Sentiment Classification
本文通过在VADER情感词典中扩展300个尼日利亚皮钦英语特有的情感词,并对14,000条皮钦英语推文进行情感标注,实现了对尼日利亚皮钦英语的语义增强。通过结合有限的人工标注数据与合成的代码混用文本,并利用专家验证的标签优化情感评分,更新后的词典在捕捉社交媒体内容中本地化语义变化方面,显著提升了原始VADER模型的上下文情感分类准确率。
Nigerian English adaptation, Pidgin, has evolved over the years through multi-language code switching, code mixing and linguistic adaptation. While Pidgin preserves many of the words in the normal English language corpus, both in spelling and pronunciation, the fundamental meaning of these words have changed significantly. For example,'ginger' is not a plant but an expression of motivation and 'tank' is not a container but an expression of gratitude. The implication is that the current approach of using direct English sentiment analysis of social media text from Nigeria is sub-optimal, as it will not be able to capture the semantic variation and contextual evolution in the contemporary meaning of these words. In practice, while many words in Nigerian Pidgin adaptation are the same as the standard English, the full English language based sentiment analysis models are not designed to capture the full intent of the Nigerian pidgin when used alone or code-mixed. By augmenting scarce human labelled code-changed text with ample synthetic code-reformatted text and meaning, we achieve significant improvements in sentiment scoring. Our research explores how to understand sentiment in an intrasentential code mixing and switching context where there has been significant word localization.This work presents a 300 VADER lexicon compatible Nigerian Pidgin sentiment tokens and their scores and a 14,000 gold standard Nigerian Pidgin tweets and their sentiments labels.
研究动机与目标
- 解决标准英文情感模型在尼日利亚皮钦英语上表现欠佳的问题,原因在于语义漂移和代码混用。
- 开发一个与VADER兼容的尼日利亚皮钦情感词典,以捕捉语境演化后的词语含义。
- 创建一个包含14,000条人工标注的尼日利亚皮钦推文的黄金标准数据集,并附带情感标签。
- 通过合成数据增强和语义增强,提升句内代码混用语境下的情感分类性能。
- 验证专家标注的情感标签是否优于基于规则的模型,在捕捉本地化皮钦表达中的细微情感差异方面。
提出的方法
- 通过将300个标准英文情感词翻译为尼日利亚皮钦,扩展了VADER词典,同时考虑了语义含义的变化。
- 生成合成的代码混用文本,以扩充有限的真实皮钦推文数据,提升模型泛化能力。
- 应用更新后的VADER词典对14,000条尼日利亚皮钦推文计算综合情感得分。
- 将原始VADER模型与更新后VADER模型的情感预测结果,与专家皮钦母语者提供的黄金标准标签进行对比。
- 使用人工标注的情感标签作为真实标签,评估并优化更新后词典的性能。
- 采用从单语到合成代码混用文本的标签迁移技术,以提升情感标注的准确性。
实验结果
研究问题
- RQ1如何扩展一个与VADER兼容的情感词典,以捕捉尼日利亚皮钦英语词语相较于标准英语的语义演变?
- RQ2在代码混用的社交媒体文本中,加入300个皮钦英语专用情感词是否能显著提升情感分类准确率?
- RQ3基于规则的模型生成的情感得分,与人工标注标签相比,在捕捉尼日利亚皮钦语境中细微情感差异方面表现如何?
- RQ4在尼日利亚皮钦等低资源、代码混用语言环境下,合成数据增强是否能有效提升情感标注性能?
- RQ5皮钦英语中哪些关键语义变化导致标准英文情感词典在尼日利亚社交媒体情感分析中失效?
主要发现
- 更新后的VADER词典,包含300个皮钦英语专用情感词,其情感分类性能显著优于原始VADER模型。
- 对于句子“som teams get black, som get purple but no one fine reach our jersey wey blue”,情感得分从-0.1154(负面)提升至0.7964(正面),与专家标注一致。
- 句子“Na to delete am”在词典更新后,从原本的中性(0.0000)被正确重新分类为负面(-0.6908),与专家标签一致。
- 句子“Why greenwood dey play nw?ole you don start”被重新分类为负面(-0.2500),而非原始的正面(0.3400),正确反映了专家判断。
- 更新后的词典在捕捉上下文细微表达方面表现更优,例如“tank”表示感激,“ginger”表示动力。
- 专家标注标签在捕捉语义漂移方面优于基于规则的情感评分,凸显了在代码混用语境中使用语言特定情感词典的必要性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。