Back to Main Conference 2018
LREC 2018main

The MADAR Arabic Dialect Corpus and Lexicon

Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)

DOI:10.63317/4wv9mx9ftkrk

Abstract

In this paper, we present two resources that were created as part of the Multi Arabic Dialect Applications and Resources (MADAR) project. The first is a large parallel corpus of 25 Arabic city dialects in the travel domain. The second is a lexicon of 1,045 concepts with an average of 45 words from 25 cities per concept. These resources are the first of their kind in terms of the breadth of their coverage and the fine location granularity. The focus on cities, as opposed to regions in studying Arabic dialects, opens new avenues to many areas of research from dialectology to dialect identification and machine translation.

Details

Paper ID
lrec2018-main-535
Pages
N/A
BibKey
bouamor-etal-2018-madar
Editor
N/A
Publisher
European Language Resources Association (ELRA)
ISSN
2522-2686
ISBN
79-10-95546-00-9
Conference
Eleventh International Conference on Language Resources and Evaluation
Location
Miyazaki, Japan
Date
7 May 2018 12 May 2018

Authors

  • HB

    Houda Bouamor

  • NH

    Nizar Habash

  • MS

    Mohammad Salameh

  • WZ

    Wajdi Zaghouani

  • OR

    Owen Rambow

  • DA

    Dana Abdulrahim

  • OO

    Ossama Obeid

  • SK

    Salam Khalifa

  • FE

    Fadhl Eryani

  • AE

    Alexander Erdmann

  • KO

    Kemal Oflazer

Links