Back to Main Conference 2014
LREC 2014main
An Arabic Twitter Corpus for Subjectivity and Sentiment Analysis
Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC 2014)
Abstract
We present a newly collected data set of 8,868 gold-standard annotated Arabic feeds. The corpus is manually labelled for subjectivity and sentiment analysis (SSA) ( = 0:816). In addition, the corpus is annotated with a variety of motivated feature-sets that have previously shown positive impact on performance. The paper highlights issues posed by twitter as a genre, such as mixture of language varieties and topic-shifts. Our next step is to extend the current corpus, using online semi-supervised learning. A first sub-corpus will be released via the ELRA repository as part of this submission.