Back to Main Conference 2022
LREC 2022main

BasqueGLUE: A Natural Language Understanding Benchmark for Basque

Proceedings of the Thirteenth International Conference on Language Resources and Evaluation (LREC 2022)

DOI:10.63317/3f53mtyzdxwr

Abstract

Natural Language Understanding (NLU) technology has improved significantly over the last few years and multitask benchmarks such as GLUE are key to evaluate this improvement in a robust and general way. These benchmarks take into account a wide and diverse set of NLU tasks that require some form of language understanding, beyond the detection of superficial, textual clues. However, they are costly to develop and language-dependent, and therefore they are only available for a small number of languages. In this paper, we present BasqueGLUE, the first NLU benchmark for Basque, a less-resourced language, which has been elaborated from previously existing datasets and following similar criteria to those used for the construction of GLUE and SuperGLUE. We also report the evaluation of two state-of-the-art language models for Basque on BasqueGLUE, thus providing a strong baseline to compare upon. BasqueGLUE is freely available under an open license.

Details

Paper ID
lrec2022-main-172
Pages
pp. 1603-1612
BibKey
urbizu-etal-2022-basqueglue
Editor
N/A
Publisher
European Language Resources Association (ELRA)
ISSN
2522-2686
ISBN
79-10-95546-38-2
Conference
Thirteenth Language Resources and Evaluation Conference
Location
Marseille, France
Date
20 June 2022 25 June 2022

Authors

  • GU

    Gorka Urbizu

  • IS

    Iñaki San Vicente

  • XS

    Xabier Saralegi

  • RA

    Rodrigo Agerri

  • AS

    Aitor Soroa

Links