Back to Main Conference 2012
LREC 2012main

DGT-TM: A freely available Translation Memory in 22 languages

Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC 2012)

DOI:10.63317/2tewbj4q8j4g

Abstract

The European Commission's (EC) Directorate General for Translation, together with the EC's Joint Research Centre, is making available a large translation memory (TM; i.e. sentences and their professionally produced translations) covering twenty-two official European Union (EU) languages and their 231 language pairs. Such a resource is typically used by translation professionals in combination with TM software to improve speed and consistency of their translations. However, this resource has also many uses for translation studies and for language technology applications, including Statistical Machine Translation (SMT), terminology extraction, Named Entity Recognition (NER), multilingual classification and clustering, and many more. In this reference paper for DGT-TM, we introduce this new resource, provide statistics regarding its size, and explain how it was produced and how to use it.

Details

Paper ID
lrec2012-main-481
Pages
pp. 454-459
BibKey
steinberger-etal-2012-dgt
Editor
N/A
Publisher
European Language Resources Association (ELRA)
ISSN
2522-2686
ISBN
978-2-9517408-7-7
Conference
Eighth International Conference on Language Resources and Evaluation
Location
Istanbul, Turkey
Date
21 May 2012 27 May 2012

Authors

  • RS

    Ralf Steinberger

  • AE

    Andreas Eisele

  • SK

    Szymon Klocek

  • SP

    Spyridon Pilos

  • PS

    Patrick Schlüter

Links