Back to Main Conference 2008
LREC 2008main
Parallel Creation of Gigaword Corpora for Medium Density Languages - an Interim Report
Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC 2008)
Abstract
For increased speed in developing gigaword language resources for medium resource density languages we integrated several FOSS tools in the HUN* toolkit. While the speed and efficiency of the resulting pipeline has surpassed our expectations, our experience in developing LDC-style resource packages for Uzbek and Kurdish makes clear that neither the data collection nor the subsequent processing stages can be fully automated.