[Paper Review] ezCoref: Towards Unifying Annotation Guidelines for Coreference Resolution
ezCoref introduces a crowdsourcing-friendly coreference annotation framework with a simplified tutorial and tool, enabling non-expert annotators to achieve >90% B3 agreement with expert annotations on re-annotated passages from seven English datasets. The method reveals inconsistencies in existing guidelines—particularly regarding generic pronouns and appositives—highlighting key phenomena that should be unified in future annotation standards.
Large-scale, high-quality corpora are critical for advancing research in coreference resolution. However, existing datasets vary in their definition of coreferences and have been collected via complex and lengthy guidelines that are curated for linguistic experts. These concerns have sparked a growing interest among researchers to curate a unified set of guidelines suitable for annotators with various backgrounds. In this work, we develop a crowdsourcing-friendly coreference annotation methodology, ezCoref, consisting of an annotation tool and an interactive tutorial. We use ezCoref to re-annotate 240 passages from seven existing English coreference datasets (spanning fiction, news, and multiple other domains) while teaching annotators only cases that are treated similarly across these datasets. Surprisingly, we find that reasonable quality annotations were already achievable (>90% agreement between the crowd and expert annotations) even without extensive training. On carefully analyzing the remaining disagreements, we identify the presence of linguistic cases that our annotators unanimously agree upon but lack unified treatments (e.g., generic pronouns, appositives) in existing datasets. We propose the research community should revisit these phenomena when curating future unified annotation guidelines.
Motivation & Objective
- To develop a scalable, crowdsourcing-friendly coreference annotation methodology that reduces reliance on expert annotators.
- To evaluate whether non-expert annotators can achieve high agreement with expert annotations using minimal linguistic training.
- To identify linguistic phenomena where crowd and expert annotations diverge, signaling gaps in current annotation guidelines.
- To promote the creation of unified, community-wide annotation guidelines by exposing inconsistencies across existing datasets.
Proposed method
- Design an interactive, tutorial-driven annotation interface that teaches only coreference cases uniformly treated across existing datasets (e.g., pronouns).
- Re-annotate 240 passages from seven diverse English coreference datasets (fiction, news, etc.) using Amazon Mechanical Turk with minimal training.
- Use a B3 F1 score to quantitatively compare crowd annotations against expert annotations, measuring agreement at the mention and cluster level.
- Conduct qualitative analysis of disagreements between crowd and expert annotations to identify linguistic phenomena lacking consistent treatment.
- Analyze inter-annotator agreement across text types to assess performance differences in familiar vs. complex texts (e.g., fiction vs. cataphoric or world-knowledge-dependent texts).
- Release the annotation tool, tutorial, and all data under an open-source license to support reproducibility and future research.
Experimental results
Research questions
- RQ1Can non-expert crowdworkers achieve high agreement with expert coreference annotations using only minimal, unified guidelines?
- RQ2Which linguistic phenomena show consistent crowd agreement but inconsistent treatment across existing expert-annotated datasets?
- RQ3How does text type (e.g., fiction, news) affect inter-annotator agreement among crowdworkers?
- RQ4What are the key discrepancies between crowd and expert annotations that reveal flaws in current annotation guidelines?
Key findings
- Non-expert annotators achieved >90% B3 F1 agreement with expert annotations on re-annotated passages, demonstrating that high-quality coreference data can be collected with minimal training.
- Crowdworkers unanimously marked generic pronouns like 'you' and 'we' as coreferent, while existing datasets inconsistently treated them—some as non-referring, others as unmarked—revealing a lack of standardization.
- Appositives and copular constructions were also consistently marked as coreferent by the crowd but inconsistently handled in expert datasets, indicating a need for unified treatment.
- Inter-annotator agreement was higher for familiar texts such as childhood stories and fiction, but dropped in texts with complex cataphora or heavy world knowledge dependencies.
- The interactive tutorial received overwhelmingly positive feedback, with one annotator calling it 'the best one I’ve ever seen'—indicating strong usability and effectiveness.
- The study identifies specific linguistic phenomena—particularly generic pronouns and appositives—as prime candidates for revisiting in future unified annotation guideline development.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.