Skip to main content
QUICK REVIEW

[论文解读] Kencorpus: A Kenyan Language Corpus of Swahili, Dholuo and Luhya for Natural Language Processing Tasks

Barack Wanjawa, Lilian Wanzare|arXiv (Cornell University)|Aug 25, 2022
Natural Language Processing Techniques被引用 4
一句话总结

Kencorpus 引入了一个公开可用的斯瓦希里语、多尔卢语和卢希亚语的多语言语料库——这三种语言均为肯尼亚的低资源语言,语料库包含 5,594 项内容(560 万词的文本,177 小时的语音)。该数据集支持下游自然语言处理任务,如词性标注、机器翻译和问答系统,概念验证系统在语音转文本任务中实现了 18.87% 的词错误率(WER),在问答任务中达到 80% 的精确匹配(EM)得分,显著推动了东非地区低资源语言处理的发展。

ABSTRACT

Indigenous African languages are categorized as under-served in Natural Language Processing. They therefore experience poor digital inclusivity and information access. The processing challenge with such languages has been how to use machine learning and deep learning models without the requisite data. The Kencorpus project intends to bridge this gap by collecting and storing text and speech data that is good enough for data-driven solutions in applications such as machine translation, question answering and transcription in multilingual communities. The Kencorpus dataset is a text and speech corpus for three languages predominantly spoken in Kenya: Swahili, Dholuo and Luhya. Data collection was done by researchers from communities, schools, media, and publishers. The Kencorpus' dataset has a collection of 5,594 items - 4,442 texts (5.6M words) and 1,152 speech files (177hrs). Based on this data, Part of Speech tagging sets for Dholuo and Luhya (50,000 and 93,000 words respectively) were developed. We developed 7,537 Question-Answer pairs for Swahili and created a text translation set of 13,400 sentences from Dholuo and Luhya into Swahili. The datasets are useful for downstream machine learning tasks such as model training and translation. We also developed two proof of concept systems: for Kiswahili speech-to-text and machine learning system for Question Answering task, with results of 18.87% word error rate and 80% Exact Match (EM) respectively. These initial results give great promise to the usability of Kencorpus to the machine learning community. Kencorpus is one of few public domain corpora for these three low resource languages and forms a basis of learning and sharing experiences for similar works especially for low resource languages.

研究动机与目标

  • 通过为肯尼亚的低资源语言创建大规模、公开可访问的语料库,解决本土非洲语言在数字领域的边缘化问题。
  • 从当地社区、学校、媒体和出版机构收集并整理高质量的文本和语音数据,以支持数据驱动的自然语言处理应用。
  • 为斯瓦希里语、多尔卢语和卢希亚语开发语言特定的资源,如词性标注、问答对和并行翻译语料集。
  • 通过 Kencorpus 语料库证明在低资源非洲语言上训练和评估自然语言处理模型的可行性。
  • 为未来在低资源自然语言处理领域,特别是在多语言非洲语境下的研究与合作奠定基础。

提出的方法

  • 通过社区协作方式进行数据收集,从当地学校、媒体机构、出版商和社区成员处获取文本和语音数据。
  • 该语料库包含 4,442 份文本条目(560 万个词)和 1,152 个音频文件(177 小时语音),涵盖斯瓦希里语、多尔卢语和卢希亚语。
  • 利用专家标注的训练数据,为多尔卢语(5 万个词)和卢希亚语(9.3 万个词)开发了词性标注集。
  • 从现有文本来源构建了包含 7,537 个斯瓦希里语问答对的问答数据集。
  • 创建了包含 13,400 个句子的并行翻译语料集,将多尔卢语和卢希亚语翻译为斯瓦希里语。
  • 实施了两个概念验证系统:一个基于自动语音识别的语音转文本系统,以及一个基于监督机器学习的问答模型。

实验结果

研究问题

  • RQ1能否有效收集并整理大规模、由社区驱动的低资源非洲语言(如斯瓦希里语、多尔卢语和卢希亚语)语料库,以支持自然语言处理应用?
  • RQ2Kencorpus 语料库在多大程度上能够支持词性标注、机器翻译和问答等下游自然语言处理任务?
  • RQ3在低资源语言的语音转文本和问答系统中,使用 Kencorpus 语料库可实现怎样的性能表现?
  • RQ4在低资源环境下,社区来源的数据在质量和多样性方面与传统数据收集方法相比如何?
  • RQ5Kencorpus 框架是否可复制到其他低资源非洲语言上,以提升数字包容性?

主要发现

  • Kencorpus 语料库包含 5,594 项内容,包括 4,442 份文本文档(560 万个词)和 1,152 个音频文件(177 小时语音),为三种主要肯尼亚语言提供了重要资源。
  • 成功为多尔卢语(5 万个词)和卢希亚语(9.3 万个词)开发了词性标注集,使低资源环境下的句法分析成为可能。
  • 构建了包含 7,537 个斯瓦希里语问答对的问答数据集,支持抽取式问答模型的训练。
  • 创建了包含 13,400 个句子的并行翻译语料集,将多尔卢语和卢希亚语翻译为斯瓦希里语,推动了多语言自然语言处理应用的发展。
  • 语音转文本系统实现了 18.87% 的词错误率(WER),证明了这些语言自动语音识别的可行性。
  • 问答模型在精确匹配(EM)得分上达到 80%,表明使用 Kencorpus 数据在抽取式问答任务中表现优异。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。