Skip to main content
QUICK REVIEW

[论文解读] Application of Lexical Features Towards Improvement of Filipino Readability Identification of Children's Literature

Joseph Marvin Imperial, Ethel Ong|arXiv (Cornell University)|Jan 22, 2021
Text Readability and Simplification参考文献 16被引用 9
一句话总结

本文通过将词汇特征(如类型-词形比率、词汇密度和外来词数量)整合到机器学习模型中,提出提升菲律宾儿童读物可读性分类。通过结合这些特征与传统的句子长度、词数等特征,作者将模型准确率提高了近5%(从42%提升至47.2%),证明词汇复杂度在菲律宾语文本可读性预测中具有显著影响。

ABSTRACT

Proper identification of grade levels of children's reading materials is an important step towards effective learning. Recent studies in readability assessment for the English domain applied modern approaches in natural language processing (NLP) such as machine learning (ML) techniques to automate the process. There is also a need to extract the correct linguistic features when modeling readability formulas. In the context of the Filipino language, limited work has been done [1, 2], especially in considering the language's lexical complexity as main features. In this paper, we explore the use of lexical features towards improving the development of readability identification of children's books written in Filipino. Results show that combining lexical features (LEX) consisting of type-token ratio, lexical density, lexical variation, foreign word count with traditional features (TRAD) used by previous works such as sentence length, average syllable length, polysyllabic words, word, sentence, and phrase counts increased the performance of readability models by almost a 5% margin (from 42% to 47.2%). Further analysis and ranking of the most important features were shown to identify which features contribute the most in terms of reading complexity.

研究动机与目标

  • 解决菲律宾语可读性评估中词汇复杂度研究不足的问题。
  • 利用自然语言处理技术,改进菲律宾儿童文学的自动化可读性识别。
  • 评估词汇特征对基于机器学习的可读性分类的影响。
  • 识别决定菲律宾语文本阅读难度复杂度的最重要特征。

提出的方法

  • 作者从菲律宾儿童读物中提取包括类型-词形比率、词汇密度、词汇多样性及外来词数量在内的词汇特征。
  • 同时收集传统可读性特征,如句子长度、平均音节数及词/短语数量。
  • 使用词汇特征(LEX)与传统特征(TRAD)的组合训练机器学习模型,以预测年级水平可读性。
  • 通过准确率指标评估模型性能,并利用模型可解释性分析对特征重要性进行排序。
  • 数据集由标注有年级水平标签的菲律宾儿童读物组成,用于模型的训练与测试。
  • 开展对比实验,比较仅使用传统特征的模型与同时包含词汇和传统特征的模型。

实验结果

研究问题

  • RQ1词汇特征在提升菲律宾儿童读物可读性分类准确率方面有何贡献?
  • RQ2哪些具体的词汇特征最能预测菲律宾语文本的阅读难度复杂度?
  • RQ3结合词汇与传统特征的模型在多大程度上优于仅使用传统特征的模型?
  • RQ4在可读性预测中,词汇特征与传统特征的重要性排名有何差异?

主要发现

  • 将词汇特征与传统特征结合后,模型准确率从42%提升至47.2%,提升近5%。
  • 类型-词形比率与词汇密度被证明是预测阅读复杂度最具影响力的词汇特征之一。
  • 外来词数量表现出显著贡献,反映出借词使用对文本难度的影响。
  • 词汇多样性与词汇密度在特征重要性分析中始终位居前列。
  • 组合模型(LEX + TRAD)显著优于仅使用传统特征的模型。
  • 结果证实,词汇复杂度是评估菲律宾儿童读物可读性的关键因素。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。