[论文解读] Bias and Agreement in Syntactic Annotations.
本文研究了人类句法标注中的锚定偏差,表明编辑器在修改解析器输出时会高估解析性能,并产生质量较低的标注。基于Penn Treebank WSJ数据,本文提供了首次受控的标注者间一致性估计,显示其一致性水平与当前最先进解析器相当,这对标注设计和解析器评估具有重要影响。
We present a study on two key characteristics of human syntactic annotations: anchoring and agreement. Anchoring is a well known cognitive bias in human decision making, where judgments are drawn towards pre-existing values. We study the influence of anchoring on a standard approach to creation of syntactic resources where syntactic annotations are obtained via human editing of tagger and parser output. Our experiments demonstrate a clear anchoring effect and reveal unwanted consequences, including overestimation of parsing performance and lower quality of annotations in comparison with human-based annotations. Using sentences from the Penn Treebank WSJ, we also report the first systematically obtained inter-annotator agreement estimates for English syntactic parsing. Our agreement results control for anchoring bias, and are consequential in that they are \emph{on par} with state of the art parsing performance for English. We discuss the impact of our findings on strategies for future annotation efforts and parser evaluations.
研究动机与目标
- 研究锚定偏差如何影响人类在编辑解析器输出时的句法标注。
- 量化锚定对标注质量和解析性能感知的影响。
- 首次系统性地提供英语句法解析的标注者间一致性估计。
- 评估锚定与一致性对今后标注策略和解析器评估的影响。
提出的方法
- 使用人类标注者编辑来自Penn Treebank WSJ语料库的解析器生成的句法结构,开展实验。
- 通过对比编辑结果与初始解析器输出,采用受控标注程序隔离并测量锚定效应。
- 在实验设计中通过控制锚定偏差,收集标注者间一致性得分。
- 使用统计分析将标注质量与性能估计值与基于人工的黄金标准标注进行比较。
- 在偏差受控条件下报告标注者间一致性度量,以确保可靠性。
- 将一致性水平与当前最先进解析性能进行对比,以 contextualize 研究发现。
实验结果
研究问题
- RQ1锚定偏差在多大程度上影响人类在编辑解析器输出时的句法标注?
- RQ2锚定如何影响标注任务中对解析性能的感知与实际表现?
- RQ3在控制锚定偏差的情况下,句法解析的真实标注者间一致性水平是多少?
- RQ4受控的一致性估计与英语中最先进解析性能相比如何?
- RQ5锚定与一致性对今后标注与评估实践有何影响?
主要发现
- 锚定偏差显著影响句法标注,导致对解析性能的高估。
- 通过编辑解析器输出生成的标注质量低于黄金标准人工标注。
- 在控制锚定偏差后,标注者间一致性达到与当前最先进解析模型相当的水平。
- 受控的一致性估计为解析系统评估提供了可靠基准。
- 锚定效应扭曲了句法标注的质量与感知准确性。
- 研究结果强调了需要修订标注方法与评估协议,以减轻偏差影响。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。