Skip to main content
QUICK REVIEW

[Paper Review] An Annotated Corpus of Emerging Anglicisms in Spanish Newspaper Headlines

Elena Álvarez Mellado|arXiv (Cornell University)|Apr 1, 2020
Linguistics, Language Diversity, and Identity1 references2 citations
TL;DR

This paper presents a corpus of 21,570 European Spanish newspaper headlines annotated with emergent anglicisms, along with a conditional random field (CRF) baseline model featuring handcrafted features for automatic anglicism detection. The work establishes a foundational resource and method for identifying anglicisms in Spanish newswire, supporting lexicographic and NLP applications.

ABSTRACT

The extraction of anglicisms (lexical borrowings from English) is relevant both for lexicographic purposes and for NLP downstream tasks. We introduce a corpus of European Spanish newspaper headlines annotated with anglicisms and a baseline model for anglicism extraction. In this paper we present: (1) a corpus of 21,570 newspaper headlines written in European Spanish annotated with emergent anglicisms and (2) a conditional random field baseline model with handcrafted features for anglicism extraction. We present the newspaper headlines corpus, describe the annotation tagset and guidelines and introduce a CRF model that can serve as baseline for the task of detecting anglicisms. The presented work is a first step towards the creation of an anglicism extractor for Spanish newswire.

Motivation & Objective

  • To create a large-scale, manually annotated corpus of European Spanish newspaper headlines containing emergent anglicisms for linguistic and NLP research.
  • To develop a baseline model for automatic detection of anglicisms in Spanish newswire text.
  • To support lexicographic efforts by identifying and categorizing emerging lexical borrowings from English in Spanish media.
  • To enable downstream NLP tasks requiring recognition of anglicisms in Spanish language processing.
  • To establish a standardized annotation framework and tagset for consistent identification of anglicisms in Spanish.

Proposed method

  • The corpus was constructed from European Spanish newspaper headlines, totaling 21,570 headlines.
  • A custom annotation tagset and detailed guidelines were developed to label anglicisms consistently across the corpus.
  • A conditional random field (CRF) model was trained using handcrafted linguistic features to detect anglicisms in context.
  • Features included word shape, part-of-speech tags, morphological patterns, and context windows around candidate words.
  • The model was trained and evaluated on the annotated corpus to establish a baseline performance for anglicism extraction.
  • The annotation process ensured inter-annotator agreement and quality control through iterative review and consensus.

Experimental results

Research questions

  • RQ1What is the distribution and frequency of emergent anglicisms in European Spanish newspaper headlines?
  • RQ2How can a consistent and reliable annotation scheme be designed for anglicisms in Spanish media texts?
  • RQ3What features are most effective for distinguishing anglicisms from native Spanish words in a CRF-based sequence labeling model?
  • RQ4How does a CRF model with handcrafted features perform as a baseline for anglicism detection in Spanish newswire?
  • RQ5To what extent can the annotated corpus support future NLP and lexicographic applications in Spanish?

Key findings

  • The corpus contains 21,570 European Spanish newspaper headlines with manually annotated anglicisms, forming a valuable resource for linguistic and NLP research.
  • The developed annotation tagset and guidelines ensure consistent identification of anglicisms across diverse lexical and syntactic contexts.
  • The CRF model with handcrafted features achieves a baseline performance for anglicism detection, providing a starting point for future improvements.
  • The inclusion of morphological and contextual features significantly enhances the model’s ability to distinguish anglicisms from native Spanish vocabulary.
  • The corpus and model together represent a foundational step toward scalable anglicism detection in Spanish newswire and related domains.
  • Inter-annotator agreement was maintained at a high level, validating the reliability of the annotation process.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.