Back to Main Conference 2012
LREC 2012main
Announcing Prague Czech-English Dependency Treebank 2.0
Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC 2012)
Abstract
We introduce a substantial update of the Prague Czech-English Dependency Treebank, a parallel corpus manually annotated at the deep syntactic layer of linguistic representation. The English part consists of the Wall Street Journal (WSJ) section of the Penn Treebank. The Czech part was translated from the English source sentence by sentence. This paper gives a high level overview of the underlying linguistic theory (the so-called tectogrammatical annotation) with some details of the most important features like valency annotation, ellipsis reconstruction or coreference.