Skip to main content
QUICK REVIEW

[论文解读] The distribution of information content in English sentences

Shuiyuan Yu, Cong Jin|arXiv (Cornell University)|Sep 24, 2016
Natural Language Processing Techniques参考文献 31被引用 4
一句话总结

本研究利用真实语料数据,通过熵及其相关统计量分析了英语句子中不同位置的信息内容分布。研究揭示出一种三步阶梯状模式:句首位置的熵低于句中位置,而句中位置又低于句尾位置,这一发现挑战了均匀信息密度(Uniform Information Density)和恒定熵率(Constant Entropy Rate)假说,表明句子中的语境信息处理具有非加和性、整体性特征。

ABSTRACT

Sentence is a basic linguistic unit, however, little is known about how information content is distributed across different positions of a sentence. Based on authentic language data of English, the present study calculated the entropy and other entropy-related statistics for different sentence positions. The statistics indicate a three-step staircase-shaped distribution pattern, with entropy in the initial position lower than the medial positions (positions other than the initial and final), the medial positions lower than the final position and the medial positions showing no significant difference. The results suggest that: (1) the hypotheses of Constant Entropy Rate and Uniform Information Density do not hold for the sentence-medial positions; (2) the context of a word in a sentence should not be simply defined as all the words preceding it in the same sentence; and (3) the contextual information content in a sentence does not accumulate incrementally but follows a pattern of "the whole is greater than the sum of parts".

研究动机与目标

  • 研究信息内容在英语句子不同位置的分布方式。
  • 检验均匀信息密度(UID)和恒定熵率(CER)假说在句子层面的有效性。
  • 考察句子中的语境信息是呈累积性增长,还是遵循非加和性、整体性模式。
  • 确定词语的语境是否应仅由其句中前序词语定义。

提出的方法

  • 使用真实语料数据,计算英语句子中句首、句中和句尾位置的熵及其相关统计量。
  • 将句子位置定义为:句首(第一个词)、句中(除首尾外的所有词)和句尾(最后一个词)以供分析。
  • 应用统计检验方法,比较三类位置的熵值差异。
  • 采用大规模真实世界英语语料,以确保研究发现的实证有效性。
  • 通过对比全句语境与逐词累积语境,分析语境信息的非加和性特征。

实验结果

研究问题

  • RQ1英语句子的句首、句中和句尾位置的信息内容如何分布?
  • RQ2在考察句子位置时,均匀信息密度假说是否成立?
  • RQ3恒定熵率假说对句子中段位置是否有效?
  • RQ4句子中的语境信息是否从其前序词语线性累积?
  • RQ5全句语境在多大程度上超过各词语语境之和?

主要发现

  • 句首位置的熵显著低于句中位置,表明句子起始部分的信息含量较低。
  • 句中位置的熵低于句尾位置,表明句尾词语承载更高的信息含量。
  • 句中各位置之间的熵无显著差异,表明句子主体部分的信息密度保持一致。
  • 研究结果反驳了句子中段位置的均匀信息密度和恒定熵率假说。
  • 语境信息含量并非线性累积,而是呈现‘整体大于部分之和’的模式,表明其处理具有非加和性特征。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。