[论文解读] Syntactic Surprisal From Neural Models Predicts, But Underestimates, Human Processing Difficulty From Syntactic Ambiguities
本研究调查了神经语言模型是否因对句法可预测性加权不足而低估了人类的花园路径效应。通过提出一种独立于词汇可预测性的句法意外度估计方法,作者发现尽管句法意外度能改善对处理难度的预测,但仍无法完全捕捉人类在歧义句中阅读时间延迟的幅度。
Humans exhibit garden path effects: When reading sentences that are temporarily structurally ambiguous, they slow down when the structure is disambiguated in favor of the less preferred alternative. Surprisal theory (Hale, 2001; Levy, 2008), a prominent explanation of this finding, proposes that these slowdowns are due to the unpredictability of each of the words that occur in these sentences. Challenging this hypothesis, van Schijndel & Linzen (2021) find that estimates of the cost of word predictability derived from language models severely underestimate the magnitude of human garden path effects. In this work, we consider whether this underestimation is due to the fact that humans weight syntactic factors in their predictions more highly than language models do. We propose a method for estimating syntactic predictability from a language model, allowing us to weigh the cost of lexical and syntactic predictability independently. We find that treating syntactic predictability independently from lexical predictability indeed results in larger estimates of garden path. At the same time, even when syntactic predictability is independently weighted, surprisal still greatly underestimate the magnitude of human garden path effects. Our results support the hypothesis that predictability is not the only factor responsible for the processing cost associated with garden path sentences.
研究动机与目标
- 调查神经语言模型对人类花园路径效应的低估是否源于对句法可预测性加权不足。
- 开发一种在神经语言模型中独立估计句法意外度的方法。
- 检验引入句法意外度是否能改善对句法歧义句中人类阅读时间延迟的预测。
- 评估模型预测与人类数据之间持续存在的差距是否意味着仅靠意外度理论无法解释花园路径效应。
提出的方法
- 提出一种利用预训练语言模型的CCG超标签估计句法意外度的方法,将句法结构与词汇内容分离。
- 通过在句法解析任务上训练的模型,计算给定上下文下一一预测句法类别的负对数概率,作为句法意外度。
- 将句法意外度与词汇意外度估计值结合,为花园路径句和无歧义句中的每个词计算综合意外度得分。
- 利用综合意外度得分预测阅读时间,并应用从非花园路径句中推导出的转换因子。
- 在三种句法结构上,将模型预测的花园路径效应大小与实证的人类阅读时间数据进行比较。
- 评估预测与观测到的阅读时间差异之间的相关性,检验句法意外度是否比仅依赖词汇意外度更能捕捉人类处理难度。
实验结果
研究问题
- RQ1将句法意外度与词汇意外度分离,是否能改善神经语言模型对人类花园路径效应的预测?
- RQ2句法意外度在多大程度上解释了歧义句中人类阅读时间延迟的幅度?
- RQ3花园路径效应的持续低估是否源于神经语言模型对句法结构加权方式的局限性?
- RQ4一种能独立处理句法与词汇可预测性的模型,是否能缩小预测与观测到的处理难度之间的差距?
主要发现
- 当句法意外度独立于词汇意外度进行估计时,其对花园路径效应幅度的预测值大于仅使用词汇意外度的结果。
- 尽管引入了句法意外度,模型预测仍显著低估了花园路径句中人类阅读时间延迟的幅度。
- 研究发现词频与句法意外度呈正相关,表明句法意外度捕捉了不依赖于词汇频率的可预测性维度。
- 识别出词汇意外但句法可预测的词语,证实句法意外度与词汇意外度并非完全相关。
- 结果表明,即使句法可预测性估计得到改进,意外度理论仍无法完全解释花园路径结构中的人类处理难度。
- 模型预测与人类数据之间持续存在的差异表明,除可预测性之外的因素——可能涉及重新分析过程——或为花园路径效应的根源。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。