Back to Main Conference 2004
LREC 2004main

Creation and Validation of Large Lexica for Speech-to-Speech Translation Purposes

Proceedings of the Fourth International Conference on Language Resources and Evaluation (LREC 2004)

DOI:10.63317/5kg45a89794r

Abstract

This paper presents specifications and requirements for creation and validation of large lexica that are needed in automatic Speech Recognition (ASR), Text-to-Speech (TTS) and statistical Speech-to-Speech Translation (SST) systems. The prepared language resources are created and validated within the scope of the EU-project LC-STAR (Lexica and Corpora for Speech-to-Speech Translation Components) during years 2002-2005. Large lexica consisting of phonetic, suprasegmental and morpho-syntactic content will be provided with well-documented specifications for 13 languages. A short summary of the LC-STAR project itself is presented. Overview about the specification for the corpora collection and word extraction as well as the specification and format of the lexica are presented. Particular attention is paid to the validation of the produced lexica and the lessons learnt during pre-validation. The created and validated language resources will be available via ELRA/ELDA.

Details

Paper ID
lrec2004-main-268
Pages
N/A
BibKey
fersoe-etal-2004-creation
Editor
N/A
Publisher
European Language Resources Association (ELRA)
ISSN
2522-2686
ISBN
2-9517408-1-6
Conference
Fourth International Conference on Language Resources and Evaluation
Location
Lisbon, Portugal
Date
26 May 2004 28 May 2004

Authors

  • HF

    Hanne Fersøe

  • EH

    Elviira Hartikainen

  • Hv

    Henk van den Heuvel

  • GM

    Giulio Maltese

  • AM

    Asuncíon Moreno

  • SS

    Shaunie Shammass

  • UZ

    Ute Ziegenhain

Links