Back to Main Conference 2008
LREC 2008main

Annotation Tool Development for Large-Scale Corpus Creation Projects at the Linguistic Data Consortium

Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC 2008)

DOI:10.63317/567oi6zrcq6o

Abstract

The Linguistic Data Consortium (LDC) creates a variety of linguistic resources - data, annotations, tools, standards and best practices - for many sponsored projects. The programming staff at LDC has created the tools and technical infrastructures to support the data creation efforts for these projects, creating tools and technical infrastructures for all aspects of data creation projects: data scouting, data collection, data selection, annotation, search, data tracking and worklow management. This paper introduces a number of samples of LDC programming staff’s work, with particular focus on the recent additions and updates to the suite of software tools developed by LDC. Tools introduced include the GScout Web Data Scouting Tool, LDC Data Selection Toolkit, ACK - Annotation Collection Kit, XTrans Transcription and Speech Annotation Tool, GALE Distillation Toolkit, and the GALE MT Post Editing Workflow Management System.

Details

Paper ID
lrec2008-main-465
Pages
N/A
BibKey
maeda-etal-2008-annotation
Editor
N/A
Publisher
European Language Resources Association (ELRA)
ISSN
2522-2686
ISBN
2-9517408-4-0
Conference
Sixth International Conference on Language Resources and Evaluation
Location
Marrakech, Morocco
Date
28 May 2008 30 May 2008

Authors

  • KM

    Kazuaki Maeda

  • HL

    Haejoong Lee

  • SM

    Shawn Medero

  • JM

    Julie Medero

  • RP

    Robert Parker

  • SS

    Stephanie Strassel

Links