Back to Main Conference 2014
LREC 2014main

ANCOR_Centre, a large free spoken French coreference corpus: description of the resource and reliability measures

Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC 2014)

DOI:10.63317/2o9k3ppc7c6e

Abstract

This article presents ANCOR_Centre, a French coreference corpus, available under the Creative Commons Licence. With a size of around 500,000 words, the corpus is large enough to serve the needs of data-driven approaches in NLP and represents one of the largest coreference resources currently available. The corpus focuses exclusively on spoken language, it aims at representing a certain variety of spoken genders. ANCOR_Centre includes anaphora as well as coreference relations which involve nominal and pronominal mentions. The paper describes into details the annotation scheme and the reliability measures computed on the resource.

Details

Paper ID
lrec2014-main-169
Pages
pp. 843-847
BibKey
muzerelle-etal-2014-ancor
Editor
N/A
Publisher
European Language Resources Association (ELRA)
ISSN
2522-2686
ISBN
978-2-9517408-8-4
Conference
Ninth International Conference on Language Resources and Evaluation
Location
Reykjavik, Iceland
Date
26 May 2014 31 May 2014

Authors

  • JM

    Judith Muzerelle

  • AL

    Anaïs Lefeuvre

  • ES

    Emmanuel Schang

  • JA

    Jean-Yves Antoine

  • AP

    Aurore Pelletier

  • DM

    Denis Maurel

  • IE

    Iris Eshkol

  • JV

    Jeanne Villaneau

Links