Skip to main content
QUICK REVIEW

[论文解读] On the Effect of Anticipation on Reading Times

Tiago Pimentel, Clara Meister|arXiv (Cornell University)|Nov 25, 2022
Text Readability and Simplification被引用 4
一句话总结

本文研究阅读时间是否受预期加工(即读者对即将出现的词语进行预测)的影响,而非仅仅由词语意外度(surprisal)所驱动的反应式加工。利用四个自然阅读数据集,研究发现上下文熵(衡量预测不确定性的一种指标)在三个数据集中比意外度更能预测阅读时间,且Rényi熵(α=1/2)优于Shannon熵,表明读者是基于可能的后续词语数量进行预期,而非基于预期的意外度。

ABSTRACT

Over the past two decades, numerous studies have demonstrated how less predictable (i.e., higher surprisal) words take more time to read. In general, these studies have implicitly assumed the reading process is purely responsive: Readers observe a new word and allocate time to process it as required. We argue that prior results are also compatible with a reading process that is at least partially anticipatory: Readers could make predictions about a future word and allocate time to process it based on their expectation. In this work, we operationalize this anticipation as a word's contextual entropy. We assess the effect of anticipation on reading by comparing how well surprisal and contextual entropy predict reading times on four naturalistic reading datasets: two self-paced and two eye-tracking. Experimentally, across datasets and analyses, we find substantial evidence for effects of contextual entropy over surprisal on a word's reading time (RT): in fact, entropy is sometimes better than surprisal in predicting a word's RT. Spillover effects, however, are generally not captured by entropy, but only by surprisal. Further, we hypothesize four cognitive mechanisms through which contextual entropy could impact RTs -- three of which we are able to design experiments to analyze. Overall, our results support a view of reading that is not just responsive, but also anticipatory.

研究动机与目标

  • 探究阅读行为是否不仅由已看到词语的反应式加工驱动,也由对即将出现词语的预期预测所驱动。
  • 检验上下文熵(衡量词语预测不确定性)是否比意外度更有效地预测阅读时间。
  • 通过四种认知机制(词语跳过、时间预算分配、预处理、不确定性成本)探讨预期如何影响阅读时间。
  • 评估Rényi熵(α=1/2)是否比Shannon熵更好地表征读者的预期。
  • 确定预期效应是否在意外度无法解释的阅读时间数据中仍可被检测到。

提出的方法

  • 将预期操作化为上下文熵,使用Rényi熵(α=1/2)作为未来词语预测不确定性的度量。
  • 在四个数据集中比较意外度与上下文熵对阅读时间的预测能力:两个自定节奏阅读数据集和两个眼动追踪数据集。
  • 使用线性混合效应模型评估意外度与熵对阅读时间的预测能力,同时控制词长及其他协变量。
  • 设计实验以检验四种预期的认知机制:词语跳过、时间预算分配、预处理和不确定性成本。
  • 应用Rényi熵的次可加性与超可加性特性,验证其作为预测因子的理论稳健性。
  • 分析数据集统计信息,并计算意外度与熵之间的Spearman等级相关系数,以排除结果由噪声导致的解释。

实验结果

研究问题

  • RQ1上下文熵是否比意外度更能预测阅读时间,从而表明存在预期加工?
  • RQ2是否存在特定的认知机制(如词语跳过或预处理)使预期影响阅读时间?
  • RQ3Rényi熵(α=1/2)是否比Shannon熵更好地操作化了预期?
  • RQ4阅读时间中的溢出效应是否源于意外度而非上下文熵?
  • RQ5熵的预测能力是否独立于意外度估计中的噪声,正如意外度与熵之间存在高相关性所暗示的那样?

主要发现

  • 在四个数据集中的三个里,上下文熵是阅读时间的显著预测因子,且在其中两个数据集中,其预测能力优于意外度。
  • Rényi熵(α=1/2)在所有情况下均显著优于Shannon熵,表明读者可能基于可能的后续词语数量进行预期。
  • 阅读时间中的溢出效应由意外度捕捉,但未被上下文熵捕捉,表明熵反映的是预期而非事后处理效应。
  • 意外度与熵之间存在强相关性(见图3),但该相关性无法解释熵的更强预测能力,从而排除了简单噪声平均作为原因的可能性。
  • 结果支持一种非纯粹反应式的阅读模型,读者会根据对后续词语的预期分配处理时间。
  • 本研究提供了实证证据,表明预期——尤其是基于不确定性的预测——在阅读时间分配中具有可测量的作用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。