Back to Main Conference 2012
LREC 2012main

Collecting and Using Comparable Corpora for Statistical Machine Translation

Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC 2012)

DOI:10.63317/2n89bqp42nan

Abstract

Lack of sufficient parallel data for many languages and domains is currently one of the major obstacles to further advancement of automated translation. The ACCURAT project is addressing this issue by researching methods how to improve machine translation systems by using comparable corpora. In this paper we present tools and techniques developed in the ACCURAT project that allow additional data needed for statistical machine translation to be extracted from comparable corpora. We present methods and tools for acquisition of comparable corpora from the Web and other sources, for evaluation of the comparability of collected corpora, for multi-level alignment of comparable corpora and for extraction of lexical and terminological data for machine translation. Finally, we present initial evaluation results on the utility of collected corpora in domain-adapted machine translation and real-life applications.

Details

Paper ID
lrec2012-main-554
Pages
pp. 438-445
BibKey
skadina-etal-2012-collecting
Editor
N/A
Publisher
European Language Resources Association (ELRA)
ISSN
2522-2686
ISBN
978-2-9517408-7-7
Conference
Eighth International Conference on Language Resources and Evaluation
Location
Istanbul, Turkey
Date
21 May 2012 27 May 2012

Authors

  • IS

    Inguna Skadiņa

  • AA

    Ahmet Aker

  • NM

    Nikos Mastropavlos

  • FS

    Fangzhong Su

  • DT

    Dan Tufis

  • MV

    Mateja Verlic

  • AV

    Andrejs Vasiļjevs

  • BB

    Bogdan Babych

  • PC

    Paul Clough

  • RG

    Robert Gaizauskas

  • NG

    Nikos Glaros

  • MP

    Monica Lestari Paramita

  • MP

    Mārcis Pinnis

Links