Skip to main content
QUICK REVIEW

[论文解读] How consistent are our discourse annotations? Insights from mapping RST-DT and PDTB annotations.

Vera Demberg, Fatemeh Torabi Asr|arXiv (Cornell University)|Apr 28, 2017
Natural Language Processing Techniques被引用 7
一句话总结

本文将宾夕法尼亚话语树库(PDTB)与修辞结构理论话语树库(RST-DT)进行映射,以评估不同话语标注框架之间的一致性。该研究提出一种基于文本片段和结构相似性的分段级对齐方法,发现显式关系的标注一致性较高,但隐式关系的标注一致性却出人意料地低,主要原因在于两个框架在标注目标和操作化定义上的差异。

ABSTRACT

Discourse-annotated corpora are an important resource for the community. However, these corpora are often annotated according to different frameworks, making comparison of the annotations difficult. This is unfortunate, since mapping the existing annotations would result in more (training) data for researchers in automatic discourse relation processing and researchers in linguistics and psycholinguistics. In this article, we present an effort to map two large corpora onto each other: the Penn Discourse Treebank and the Rhetorical Structure Theory Discourse Treebank. We first propose a method for aligning the discourse segments, and then evaluate the observed against the expected mappings for explicit and implicit relations separately. We find that while agreement on explicit relations is reasonable, agreement between the frameworks on implicit relations is astonishingly low. We identify sources of systematic discrepancies between the two annotation schemes; many of the differences in annotation can be traced back to different operationalizations and goals of the PDTB and RST frameworks. We discuss the consequences of these discrepancies for future annotation, and the usability of the mapped data for theoretical studies and the training of automatic discourse relation labellers.

研究动机与目标

  • 评估两大主要话语语料库PDTB与RST-DT之间话语标注的一致性。
  • 解决在不同理论框架之间比较话语标注的挑战,该挑战阻碍了数据整合与模型训练。
  • 识别PDTB与RST-DT标注方案之间的系统性差异,特别是针对隐式话语关系。
  • 评估映射后数据在理论话语研究与自动话语关系标注中的可用性。
  • 通过分析不同框架目标与操作化方式的影响,为未来话语标注实践提供建议。

提出的方法

  • 提出一种基于文本跨度和结构相似性的方法,用于对齐PDTB与RST-DT中的话语片段。
  • 通过将实际映射结果与预期映射结果进行对比,评估对齐效果,分别针对显式与隐式话语关系进行分析。
  • 通过考察两个框架在标注目标和操作定义上的差异,分析其标注方案之间的差异。
  • 使用定量评估指标,分别测量显式与隐式关系的一致性。
  • 通过分析每个框架如何定义和应用话语关系类别,识别不一致的根源。
  • 利用映射后的数据评估其在训练自动话语关系分类器及支持理论研究方面的潜力。

实验结果

研究问题

  • RQ1在分段级别对齐后,PDTB与RST-DT的话语标注一致性如何?
  • RQ2PDTB与RST-DT在显式话语关系上的共识程度与隐式关系相比如何?
  • RQ3PDTB与RST-DT标注方案之间系统性差异的主要来源是什么?
  • RQ4两个框架在操作化方式与目标上的差异如何影响标注一致性?
  • RQ5映射后的数据在多大程度上可用于训练自动话语关系标注器或支持理论话语研究?

主要发现

  • PDTB与RST-DT在显式话语关系上的标注一致性合理较高,表明这些情况下的标注具有一致性。
  • PDTB与RST-DT在隐式话语关系上的标注一致性出人意料地低,表明这些关系在不同框架中的标注存在根本性不一致。
  • 标注中的系统性差异主要归因于PDTB与RST框架在操作定义与理论目标上的差异。
  • 隐式关系标注一致性低,削弱了将两个语料库数据合并用于训练自动话语关系分类器的可靠性。
  • 映射数据表明,标注差异并非随机产生,而是源于不同的设计原则,从而限制了跨框架的可比性。
  • 研究结果强调了提升话语标注框架之间一致性的重要性,以增强数据可重用性与理论一致性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。