HomeLREC 2022WorkshopsSIGULlrec2022-ws-sigul-07
Back to SIGUL 2022
LREC 2022workshop

Tupían Language Ressources: Data, Tools, Analyses

Proceedings of the 1st Annual Meeting of the ELRA/ISCA Special Interest Group on Under-Resourced Languages

DOI:10.63317/4nk7outs3ikt

Abstract

TuLaR (Tupian Language Resources) is a project for collecting, documenting, analyzing, and developing computational and pedagogical material for low-resource Brazilian indigenous languages. It provides valuable data for language research regarding typological, syntactic, morphological, and phonological aspects. Here we present TuLaR’s databases, with special consideration to TuDeT (Tupian Dependency Treebanks), an annotated corpus under development for nine languages of the Tupian family, built upon the Universal Dependencies framework. The annotation within such a framework serves a twofold goal: enriching the linguistic documentation of the Tupian languages due to the rapid and consistent annotation, and providing computational resources for those languages, thanks to the suitability of our framework for developing NLP tools. We likewise present a related lexical database, some tools developed by the project, and examine future goals for our initiative.

Details

Paper ID
lrec2022-ws-sigul-07
Pages
pp. 48-58
BibKey
martin-rodriguez-etal-2022-tupian
Editor
N/A
Publisher
European Language Resources Association (ELRA)
ISSN
N/A
ISBN
N/A
Workshop
Proceedings of the 1st Annual Meeting of the ELRA/ISCA Special Interest Group on Under-Resourced Languages
Location
undefined, undefined
Date
20 June 2022 25 June 2022

Authors

  • LM

    Lorena Martín Rodríguez

  • TM

    Tatiana Merzhevich

  • WS

    Wellington Silva

  • TT

    Tiago Tresoldi

  • CA

    Carolina Aragon

  • FG

    Fabrício F. Gerardi

Links