Skip to main content
QUICK REVIEW

[论文解读] Implicit Knowledge in Argumentative Texts: An Annotated Corpus

Maria Becker, Katharina Korfhage|arXiv (Cornell University)|Dec 4, 2019
Topic Modeling参考文献 16被引用 6
一句话总结

本文提出了一项高质量、标注丰富的论说文中隐含知识语料库,其中人工标注者为来自 Microtexts 语料库的论说性单元补全缺失前提,以建立其间的逻辑连接。通过进一步将这些插入内容标注为语义子句类型及基于 ConceptNet 的常识关系,研究发现通用陈述和因果关系(如 Causes)在填补论说性空白中占主导地位,为自动化论说分析与隐含知识重建提供了关键洞见。

ABSTRACT

When speaking or writing, people omit information that seems clear and evident, such that only part of the message is expressed in words. Especially in argumentative texts it is very common that (important) parts of the argument are implied and omitted. We hypothesize that for argument analysis it will be beneficial to reconstruct this implied information. As a starting point for filling such knowledge gaps, we build a corpus consisting of high-quality human annotations of missing and implied information in argumentative texts. To learn more about the characteristics of both the argumentative texts and the added information, we further annotate the data with semantic clause types and commonsense knowledge relations. The outcome of our work is a carefully de-signed and richly annotated dataset, for which we then provide an in-depth analysis by investigating characteristic distributions and correlations of the assigned labels. We reveal interesting patterns and intersections between the annotation categories and properties of our dataset, which enable insights into the characteristics of both argumentative texts and implicit knowledge in terms of structural features and semantic information. The results of our analysis can help to assist automated argument analysis and can guide the process of revealing implicit information in argumentative texts automatically.

研究动机与目标

  • 为解决自然语言中论说不完整的问题,尤其是论说性文本中关键前提被省略或隐含的情况。
  • 构建一个高质量、人工标注的数据集,以捕捉连接论说性单元的缺失或隐含知识。
  • 利用语义子句类型和常识知识关系,分析原始论说性单元及插入的隐含知识在结构和语义上的特性。
  • 识别论说结构、语义特征与知识类型之间的模式与相关性,以指导自动化隐含知识恢复。
  • 将该数据集作为 Microtext 语料库的扩展发布,以支持未来在论说挖掘与自然语言理解方面的研究。

提出的方法

  • 标注者被提供来自 Microtext 语料库的论说性单元对,并被要求生成简短、自然语言的句子,以明确表达两者之间的隐含联系。
  • 每条插入的句子均标注语义子句类型(如 Generic、State、Event),以捕捉隐含知识的结构与泛化特征。
  • 每条插入句子还与 ConceptNet 知识关系(如 IsA、Causes、PartOf)关联,以将隐含知识的语义内容映射到常识知识结构中。
  • 对数据集进行标签分布、共现模式及论说关系、语义子句类型与常识关系之间的相关性分析。
  • 采用统计相关性分析(使用 MCC)来检验原始微文本与插入句子中语义子句类型与常识关系之间的关系。
  • 比较不同论说关系(如 support、undercut)下插入句子的特征,以评估隐含知识在结构上的依赖性。

实验结果

研究问题

  • RQ1在论说性单元之间插入的隐含知识中,哪些语义子句类型(如 Generic、State、Event)最为普遍?
  • RQ2常识知识关系(如 Causes、IsA、PartOf)在原始论说性单元与插入的隐含句子中如何分布?
  • RQ3论说关系(如 support、undercut)与插入隐含知识属性(如子句类型、知识关系)之间存在何种相关性?
  • RQ4隐含知识的结构与语义特征如何因单元间论说关系类型的不同而有所差异?
  • RQ5特定知识类型(如通用陈述、因果关系)与需要多个插入前提的需求之间相关性有多大?

主要发现

  • 通用句子在插入的隐含知识中占主导地位,表明泛化在填补论说空白中起着核心作用。
  • Causes 关系与插入句子中的 States 存在强烈正相关,表明因果解释常被用于连接支持性论说单元。
  • 在微文本与插入句子中,Generic 句子与 IsA、AtLocation、PartOf 等关系均表现出显著负相关,表明通用知识与个体特定或空间关系截然不同。
  • 复杂论说关系(如 undercut)比简单关系需要更多插入句子,表明解决此类连接的认知与信息负荷更高。
  • 表达 States 的句子在相邻论说性单元之间使用最频繁,而 Events 更常用于连接非相邻但论说相关的单元。
  • 在微文本与插入句子中,IsA 与 States 之间的相关性始终较高,表明身份与属性关系常被用于描述特定个体或事件。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。