Back to Main Conference 2022
LREC 2022main

MAKED: Multi-lingual Automatic Keyword Extraction Dataset

Proceedings of the Thirteenth International Conference on Language Resources and Evaluation (LREC 2022)

DOI:10.63317/4jzz9dazek37

Abstract

Keyword extraction is an integral task for many downstream problems like clustering, recommendation, search and classification. Development and evaluation of keyword extraction techniques require an exhaustive dataset; however, currently, the community lacks large-scale multi-lingual datasets. In this paper, we present MAKED, a large-scale multi-lingual keyword extraction dataset comprising of 540K+ news articles from British Broadcasting Corporation News (BBC News) spanning 20 languages. It is the first keyword extraction dataset for 11 of these 20 languages. The quality of the dataset is examined by experimentation with several baselines. We believe that the proposed dataset will help advance the field of automatic keyword extraction given its size, diversity in terms of languages used, topics covered and time periods as well as its focus on under-studied languages.

Details

Paper ID
lrec2022-main-664
Pages
pp. 6170-6179
BibKey
verma-etal-2022-maked
Editor
N/A
Publisher
European Language Resources Association (ELRA)
ISSN
2522-2686
ISBN
79-10-95546-38-2
Conference
Thirteenth Language Resources and Evaluation Conference
Location
Marseille, France
Date
20 June 2022 25 June 2022

Authors

  • YV

    Yash Verma

  • AJ

    Anubhav Jangra

  • SS

    Sriparna Saha

  • AJ

    Adam Jatowt

  • DR

    Dwaipayan Roy

Links