Back to Main Conference 2024
LREC-COLING 2024main

Connecting Language Technologies with Rich, Diverse Data Sources Covering Thousands of Languages

Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)

DOI:10.63317/2d9yph73q269

Abstract

Contrary to common belief, there are rich and diverse data sources available for many thousands of languages, which can be used to develop technologies for these languages. In this paper, we provide an overview of some of the major online data sources, the types of data that they provide access to, potential applications of this data, and the number of languages that they cover. Even this covers only a small fraction of the data that exists; for example, printed books are published in many languages but few online aggregators exist.

Details

Paper ID
lrec2024-main-0331
Pages
pp. 3729-3746
BibKey
van-esch-etal-2024-connecting
Editor
N/A
Publisher
European Language Resources Association (ELRA) and ICCL
ISSN
2522-2686
ISBN
979-10-95546-34-4
Conference
Joint International Conference on Computational Linguistics, Language Resources and Evaluation
Location
Turin, Italy
Date
20 May 2024 25 May 2024

Authors

  • Dv

    Daan van Esch

  • SR

    Sandy Ritchie

  • SR

    Sebastian Ruder

  • JK

    Julia Kreutzer

  • CR

    Clara Rivera

  • IS

    Ishank Saxena

  • IC

    Isaac Caswell

Links