A Twitter Corpus for Named Entity Recognition in Turkish

Proceedings of the Thirteenth International Conference on Language Resources and Evaluation (LREC 2022)

Abstract

This paper introduces a new Turkish Twitter Named Entity Recognition dataset. The dataset, which consists of 5000 tweets from a year-long period, was labeled by multiple annotators with a high agreement score. The dataset is also diverse in terms of the named entity types as it contains not only person, organization, and location but also time, money, product, and tv-show categories. Our initial experiments with pretrained language models (like BertTurk) over this dataset returned F1 scores of around 80%. We share this dataset publicly.

Resources

Details

Paper ID

lrec2022-main-484

Pages

pp. 4546-4551

DOI

10.63317/4at3je6ua8zo

BibKey

carik-yeniterzi-2022-twitter

Editors

Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Jan Odijk, Stelios Piperidis2020

Publisher

European Language Resources Association (ELRA)

ISSN

2522-2686

ISBN

79-10-95546-38-2

Conference

Thirteenth Language Resources and Evaluation Conference

Location

Marseille, France

Date

20 - 25 June 2022

Authors

BÇ
Buse Çarık
RY
Reyyan Yeniterzi

Links

URL

DOI