[Paper Review] Bias and Agreement in Syntactic Annotations.
This paper investigates anchoring bias in human syntactic annotation, demonstrating that editors overestimate parsing performance and produce lower-quality annotations when modifying parser outputs. Using Penn Treebank WSJ data, it provides the first controlled inter-annotator agreement estimates, showing agreement levels on par with state-of-the-art parsers, which has critical implications for annotation design and parser evaluation.
We present a study on two key characteristics of human syntactic annotations: anchoring and agreement. Anchoring is a well known cognitive bias in human decision making, where judgments are drawn towards pre-existing values. We study the influence of anchoring on a standard approach to creation of syntactic resources where syntactic annotations are obtained via human editing of tagger and parser output. Our experiments demonstrate a clear anchoring effect and reveal unwanted consequences, including overestimation of parsing performance and lower quality of annotations in comparison with human-based annotations. Using sentences from the Penn Treebank WSJ, we also report the first systematically obtained inter-annotator agreement estimates for English syntactic parsing. Our agreement results control for anchoring bias, and are consequential in that they are \emph{on par} with state of the art parsing performance for English. We discuss the impact of our findings on strategies for future annotation efforts and parser evaluations.
Motivation & Objective
- To investigate how anchoring bias affects human syntactic annotation when editing parser outputs.
- To quantify the impact of anchoring on annotation quality and perceived parsing performance.
- To provide the first systematically controlled estimates of inter-annotator agreement for English syntactic parsing.
- To assess the implications of anchoring and agreement for future annotation strategies and parser evaluation.
Proposed method
- Conducted experiments using human annotators to edit parser-generated syntactic structures from the Penn Treebank WSJ corpus.
- Applied controlled annotation procedures to isolate and measure anchoring effects by comparing edits to initial parser outputs.
- Collected inter-annotator agreement scores while accounting for anchoring bias through experimental design.
- Used statistical analysis to compare annotation quality and performance estimates against human-based gold-standard annotations.
- Reported inter-annotator agreement metrics under bias-controlled conditions to ensure reliability.
- Evaluated agreement levels against state-of-the-art parsing performance to contextualize findings.
Experimental results
Research questions
- RQ1To what extent does anchoring bias influence human syntactic annotation when editing parser outputs?
- RQ2How does anchoring affect perceived and actual parsing performance in annotation tasks?
- RQ3What is the true level of inter-annotator agreement in syntactic parsing when anchoring bias is controlled?
- RQ4How do controlled agreement estimates compare to state-of-the-art parsing performance in English?
- RQ5What are the implications of anchoring and agreement for future annotation and evaluation practices?
Key findings
- Anchoring bias significantly influences syntactic annotations, leading to overestimation of parsing performance.
- Annotations produced via editing parser outputs are of lower quality than gold-standard human annotations.
- Inter-annotator agreement, when controlled for anchoring, reaches levels on par with state-of-the-art parsing models.
- The controlled agreement estimates provide a reliable benchmark for evaluating parsing systems.
- Anchoring effects distort both the quality and perceived accuracy of syntactic annotations.
- The findings highlight the need for revised annotation methodologies and evaluation protocols to mitigate bias.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.