Back to Main Conference 2012
LREC 2012main

Development of a Web-Scale Chinese Word N-gram Corpus with Parts of Speech Information

Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC 2012)

DOI:10.63317/39s9e6src3br

Abstract

Web provides a large-scale corpus for researchers to study the language usages in real world. Developing a web-scale corpus needs not only a lot of computation resources, but also great efforts to handle the large variations in the web texts, such as character encoding in processing Chinese web texts. In this paper, we aim to develop a web-scale Chinese word N-gram corpus with parts of speech information called NTU PN-Gram corpus using the ClueWeb09 dataset. We focus on the character encoding and some Chinese-specific issues. The statistics about the dataset is reported. We will make the resulting corpus a public available resource to boost the Chinese language processing.

Details

Paper ID
lrec2012-main-340
Pages
pp. 320-324
BibKey
yu-etal-2012-development
Editor
N/A
Publisher
European Language Resources Association (ELRA)
ISSN
2522-2686
ISBN
978-2-9517408-7-7
Conference
Eighth International Conference on Language Resources and Evaluation
Location
Istanbul, Turkey
Date
21 May 2012 27 May 2012

Authors

  • CY

    Chi-Hsin Yu

  • YT

    Yi-jie Tang

  • HC

    Hsin-Hsi Chen

Links