Skip to main content
QUICK REVIEW

[Paper Review] BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions

Nayla Escribano, Jon Ander González|arXiv (Cornell University)|May 3, 2022
Basque language and culture studies4 citations
TL;DR

This paper introduces BasqueParl, a 14M-word publicly available bilingual corpus of Basque and Spanish parliamentary transcripts from 2012–2020, enriched with metadata (speaker, gender, party, birth year), lemmatization, and named entity recognition. The study reveals that while Basque is frequently used in speeches, it primarily serves rhetorical or formal functions rather than conveying core content, and although women are underrepresented in speech frequency and length, their word production has significantly increased since 2012, indicating a growing presence in parliamentary discourse.

ABSTRACT

Parliamentary transcripts provide a valuable resource to understand the reality and know about the most important facts that occur over time in our societies. Furthermore, the political debates captured in these transcripts facilitate research on political discourse from a computational social science perspective. In this paper we release the first version of a newly compiled corpus from Basque parliamentary transcripts. The corpus is characterized by heavy Basque-Spanish code-switching, and represents an interesting resource to study political discourse in contrasting languages such as Basque and Spanish. We enrich the corpus with metadata related to relevant attributes of the speakers and speeches (language, gender, party...) and process the text to obtain named entities and lemmas. The obtained metadata is then used to perform a detailed corpus analysis which provides interesting insights about the language use of the Basque political representatives across time, parties and gender.

Motivation & Objective

  • To create a publicly available, large-scale bilingual corpus of Basque and Spanish parliamentary transcripts to support computational analysis of political discourse.
  • To investigate language use patterns, particularly code-switching, in Basque political discourse across time, gender, and political parties.
  • To analyze gender disparities in parliamentary speech frequency and length, assessing representation and linguistic participation.
  • To enrich the corpus with structured metadata (speaker, gender, party, birth year), lemmas, and named entities for advanced NLP and social science analysis.
  • To provide empirical insights into how linguistic behavior in the Basque Parliament reflects or diverges from societal language use and political representation.

Proposed method

  • Gathered and processed transcriptions from two legislative terms (2012–2016 and 2016–2020) of the Basque Parliament.
  • Automatically identified and labeled the language of each speech fragment to capture Basque-Spanish code-switching patterns.
  • Annotated the corpus with speaker metadata: gender, party, year of birth, and legislative term.
  • Applied neural lemmatization models for both Basque and Spanish to standardize word forms.
  • Performed named entity recognition (NER) for both languages to extract entities such as institutions, people, and organizations.
  • Conducted statistical and longitudinal analysis on linguistic features, gender representation, and party-specific speech patterns.

Experimental results

Research questions

  • RQ1How does the use of Basque and Spanish vary across time, political parties, and speaker demographics in parliamentary debates?
  • RQ2To what extent is Basque used to convey speech content versus serving as a formal or rhetorical device?
  • RQ3What are the gender disparities in speech frequency and word count among Basque political representatives, and how have they evolved over time?
  • RQ4How do speech patterns differ between parties in terms of language use and gender representation?
  • RQ5To what extent does the linguistic behavior of Basque MPs reflect or diverge from broader sociolinguistic trends in the Basque Country?

Key findings

  • Basque is frequently used in parliamentary speeches but primarily for formal openings, closings, and addressing colleagues, rather than conveying core speech content.
  • Spanish is the dominant language for conveying speech content, with the most frequently mentioned entity in Spanish (Euskadi) appearing nearly four times more often than the most frequent Basque entity.
  • Women produce fewer speeches and fewer words than men overall, with a gap exceeding 13 percentage points in word production even after excluding the president’s speeches.
  • Despite lower representation in speech frequency, women’s word production has increased significantly over time, rising from less than one-third in 2012 to nearly two-thirds by 2020.
  • The party EH Bildu stands out as an exception, where female speakers show higher linguistic participation than expected based on their representation in parliament.
  • Two parties (PSE-EE and PP) exhibit higher female word production than expected from their speaker representation, while EAJ-PNV shows a 10-point deficit in female word share.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.