Skip to main content
QUICK REVIEW

[论文解读] NILC-Metrix: assessing the complexity of written and spoken language in Brazilian Portuguese

Sidney Evaldo Leal, Magali Sanches Duran|arXiv (Cornell University)|Dec 17, 2021
Text Readability and Simplification被引用 7
一句话总结

本论文介绍了NILC-Metrix,这是一个公开可用的计算系统,包含200项针对葡萄牙语(巴西)文本复杂性的语言学度量指标,其灵感来源于语篇分析、心理语言学和计算语言学。通过三个应用展示了该框架的实用性——比较儿童字幕与校本文本、预测PorSimples语料库中的复杂性,以及建模学校年级预测,展示了其在多种语言层次和任务上的稳健性。

ABSTRACT

This paper presents and makes publicly available the NILC-Metrix, a computational system comprising 200 metrics proposed in studies on discourse, psycholinguistics, cognitive and computational linguistics, to assess textual complexity in Brazilian Portuguese (BP). These metrics are relevant for descriptive analysis and the creation of computational models and can be used to extract information from various linguistic levels of written and spoken language. The metrics in NILC-Metrix were developed during the last 13 years, starting in 2008 with Coh-Metrix-Port, a tool developed within the scope of the PorSimples project. Coh-Metrix-Port adapted some metrics to BP from the Coh-Metrix tool that computes metrics related to cohesion and coherence of texts in English. After the end of PorSimples in 2010, new metrics were added to the initial 48 metrics of Coh-Metrix-Port. Given the large number of metrics, we present them following an organisation similar to the metrics of Coh-Metrix v3.0 to facilitate comparisons made with metrics in Portuguese and English. In this paper, we illustrate the potential of NILC-Metrix by presenting three applications: (i) a descriptive analysis of the differences between children's film subtitles and texts written for Elementary School I and II (Final Years); (ii) a new predictor of textual complexity for the corpus of original and simplified texts of the PorSimples project; (iii) a complexity prediction model for school grades, using transcripts of children's story narratives told by teenagers. For each application, we evaluate which groups of metrics are more discriminative, showing their contribution for each task.

研究动机与目标

  • 开发一个全面、公开可访问的葡萄牙语(巴西)文本复杂性评估语言学度量系统。
  • 将基于英语的工具(如Coh-Metrix)中的度量指标进行扩展与适配,以适用于葡萄牙语(巴西)语境,确保跨语言可比性。
  • 支持对书面与口语语言在多种语言层次上的描述性分析与计算建模。
  • 在真实世界NLP应用中验证度量套件的实用性,包括文本简化与教育文本复杂性预测。
  • 通过发布代码、数据集以及模块化框架,为未来研究提供支持,便于度量的实现与扩展。

提出的方法

  • 改编并重新实现最初在PorSimples项目及后续研究中开发的200项度量指标,涵盖衔接、连贯性、词汇、句法和语义特征。
  • 按照Coh-Metrix v3.0的结构组织度量指标,以确保与现有英语工具的一致性与可比性。
  • 使用NLP工具(包括LX-Parser、MaltParser和Palavras)进行句法解析,采用基于LSA的方法计算语义衔接度量指标。
  • 在三个不同任务中应用传统机器学习分类器,利用度量指标特征预测文本复杂性。
  • 实现新度量指标(如思想密度和词汇紧密度),并计划通过POeTiSA项目中的鲁棒解析模型进行集成。
  • 评估不同度量组在不同类型文本与复杂性水平下的判别能力,包括儿童叙事和简化文本。

实验结果

研究问题

  • RQ1在葡萄牙语(巴西)中,哪些语言学度量组最能有效区分儿童电影字幕与校本书面文本?
  • RQ2NILC-Metrix度量指标在PorSimples语料库原始版本与简化版本的文本复杂性预测中表现如何?
  • RQ3基于NILC-Metrix度量指标的模型能否成功从青少年口语叙事中预测学校年级水平?
  • RQ4不同语言层次(词汇、句法、衔接、语义)在多种文本类型中的复杂性判别中分别起到多大程度的贡献?
  • RQ5鲁棒解析模型在多大程度上能提升如思想密度和时间衔接等度量指标在葡萄牙语(巴西)中的可靠性与泛化能力?

主要发现

  • NILC-Metrix系统成功捕捉了儿童电影字幕与校本文本之间复杂性的差异,其中词汇与衔接度量指标表现出强大的判别能力。
  • 针对PorSimples语料库的文本复杂性新预测器,通过结合句法与词汇度量指标,实现了高准确率,证明了该系统在文本简化评估中的实用性。
  • 基于青少年口语叙事转录本的学校年级复杂性预测模型表现出显著性能,其中语义与衔接度量指标尤为具信息量。
  • 与词汇多样性、句子长度以及衔接度量(尤其是连接词与内容词关联)相关的度量指标,在所有三个应用中均持续表现出最强的判别力。
  • 预计未来版本中,通过集成POeTiSA项目中的鲁棒解析模型,将进一步提升思想密度与时间衔接等度量指标的可靠性。
  • 代码与数据集以AGPLv3许可证公开发布,确保了可复现性,并推动了葡萄牙语(巴西)计算语言学与NLP领域的进一步研究。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。