[论文解读] Feature overwriting as a finite mixture process: Evidence from comprehension data
该论文提出,像 'The key to the cabinets are on the table' 这类语法错误的句子之所以会产生语法正确的错觉,是因为特征重写(feature overwriting)所致,即在部分试验中,具有相同数范畴的名词被混淆。基于10个已发表的数据集,采用分层贝叶斯有限混合模型分析,发现基于特征重写的异方差混合模型在解释阅读时间加快方面,显著优于特征传播(feature percolation)和基于线索的检索(cue-based retrieval)模型。
The ungrammatical sentence "The key to the cabinets are on the table" is known to lead to an illusion of grammaticality. As discussed in the meta-analysis by Jaeger et al., 2017, faster reading times are observed at the verb are in the agreement-attraction sentence above compared to the equally ungrammatical sentence "The key to the cabinet are on the table". One explanation for this facilitation effect is the feature percolation account: the plural feature on cabinets percolates up to the head noun key, leading to the illusion. An alternative account is in terms of cue-based retrieval (Lewis & Vasishth, 2005), which assumes that the non-subject noun cabinets is misretrieved due to a partial feature-match when a dependency completion process at the auxiliary initiates a memory access for a subject with plural marking. We present evidence for yet another explanation for the observed facilitation. Because the second sentence has two nouns with identical number, it is possible that these are, in some proportion of trials, more difficult to keep distinct, leading to slower reading times at the verb in the first sentence above; this is the feature overwriting account of Nairne, 1990. We show that the feature overwriting proposal can be implemented as a finite mixture process. We reanalysed ten published data-sets, fitting hierarchical Bayesian mixture models to these data assuming a two-mixture distribution. We show that in nine out of the ten studies, a mixture distribution corresponding to feature overwriting furnishes a superior fit over both the feature percolation and the cue-based retrieval accounts.
研究动机与目标
- 探究特征重写(即具有相同数范畴的名词之间的混淆)是否能够解释语法错误句子中语法正确的错觉。
- 比较三种竞争性解释模型(特征重写、特征传播、基于线索的检索)的预测性能。
- 在分层贝叶斯框架下,使用有限混合分布对阅读时间加快背后的认知过程进行建模。
- 确定特征重写是否相比现有解释模型,对已发表的语 comprehension 数据提供了更优的统计拟合。
提出的方法
- 将分层贝叶斯双混合模型拟合至10个已发表的阅读时间数据集,假设存在两种分布的有限混合。
- 采用异方差混合模型,其中一种成分代表因特征重写导致高混淆度的试验。
- 通过一个概率参数(diffprob)定义混合模型,表示高混淆度成分中试验的比例。
- 使用PSIS-LOO交叉验证比较模型拟合,更高的elpd值表示更好的预测性能。
- 将阅读时间建模为对数正态分布,包含被试和项目层面的随机效应,并引入对实验条件的求和编码预测变量。
- 使用Stan进行完整贝叶斯推断,以估计模型参数及其不确定性。
实验结果
研究问题
- RQ1与特征传播或基于线索的检索相比,特征重写是否能为语法错误句子中动词位置的阅读时间加快提供更优的统计解释?
- RQ2是否存在证据表明,因特征重写导致高混淆度的试验,其阅读时间方差高于其他试验?
- RQ3在两个名词均为单数与一个单数一个复数名词的条件下,归因于高混淆度成分的试验比例是否存在差异?
- RQ4基于特征重写的有限混合模型能否解释语 comprehension 数据中观察到的促进效应?
主要发现
- 在十个数据集中的九个中,异方差特征重写混合模型的拟合效果优于特征传播模型和基于线索的检索模型。
- 同方差特征重写模型在除一个数据集外的所有数据集中均优于标准分层模型(即检索干扰解释),表明对特征重写的更强支持。
- 所有研究中,高混淆度成分中试验比例的估计值(diffprob)均大于零,仅在研究1中存在较高不确定性,表明在双单数名词条件下存在一致的混淆度增加证据。
- 高混淆度分布的方差(sigmap_e)显著高于其他方差分量,表明在发生特征重写时,试验间的变异性更大。
- 模型比较结果具有传递性,且异方差特征重写模型在整体预测性能上最优,如PSIS-LOO比较中更高的elpd值所示。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。