Back to Main Conference 2022
LREC 2022main

BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions

Proceedings of the Thirteenth International Conference on Language Resources and Evaluation (LREC 2022)

DOI:10.63317/4eso6vw52j7w

Abstract

Parliamentary transcripts provide a valuable resource to understand the reality and know about the most important facts that occur over time in our societies. Furthermore, the political debates captured in these transcripts facilitate research on political discourse from a computational social science perspective. In this paper we release the first version of a newly compiled corpus from Basque parliamentary transcripts. The corpus is characterized by heavy Basque-Spanish code-switching, and represents an interesting resource to study political discourse in contrasting languages such as Basque and Spanish. We enrich the corpus with metadata related to relevant attributes of the speakers and speeches (language, gender, party...) and process the text to obtain named entities and lemmas. The obtained metadata is then used to perform a detailed corpus analysis which provides interesting insights about the language use of the Basque political representatives across time, parties and gender.

Details

Paper ID
lrec2022-main-361
Pages
pp. 3382-3390
BibKey
escribano-etal-2022-basqueparl
Editor
N/A
Publisher
European Language Resources Association (ELRA)
ISSN
2522-2686
ISBN
79-10-95546-38-2
Conference
Thirteenth Language Resources and Evaluation Conference
Location
Marseille, France
Date
20 June 2022 25 June 2022

Authors

  • NE

    Nayla Escribano

  • JG

    Jon Ander Gonzalez

  • JO

    Julen Orbegozo-Terradillos

  • AL

    Ainara Larrondo-Ureta

  • SP

    Simón Peña-Fernández

  • OP

    Olatz Perez-de-Viñaspre

  • RA

    Rodrigo Agerri

Links