Back to Main Conference 2024
LREC-COLING 2024main

The ParCoLab Parallel Corpus and Its Extension to Four Regional Languages of France

Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)

DOI:10.63317/4j9icygpptfj

Abstract

Parallel corpora are still scarce for most of the world’s language pairs. The situation is by no means different for regional languages of France. In addition, adequate web interfaces facilitate and encourage the use of parallel corpora by target users, such as language learners and teachers, as well as linguists. In this paper, we describe ParCoLab, a parallel corpus and a web platform for querying the corpus. From its onset, ParCoLab has been geared towards lower-resource languages, with an initial corpus in Serbian, along with French and English (later Spanish). We focus here on the extension of ParCoLab with a parallel corpus for four regional languages of France: Alsatian, Corsican, Occitan and Poitevin-Saintongeais. In particular, we detail criteria for choosing texts and issues related to their collection. The new parallel corpus contains more than 20k tokens per regional language.

Details

Paper ID
lrec2024-main-1392
Pages
pp. 16014-16023
BibKey
stosic-etal-2024-parcolab
Editor
N/A
Publisher
European Language Resources Association (ELRA) and ICCL
ISSN
2522-2686
ISBN
979-10-95546-34-4
Conference
Joint International Conference on Computational Linguistics, Language Resources and Evaluation
Location
Turin, Italy
Date
20 May 2024 25 May 2024

Authors

  • DS

    Dejan Stosic

  • SM

    Saša Marjanović

  • DB

    Delphine Bernhard

  • XB

    Xavier Bach

  • MB

    Myriam Bras

  • LK

    Laurent Kevers

  • SR

    Stella Retali-Medori

  • MV

    Marianne Vergez-Couret

  • CW

    Carole Werner

Links