Back to Main Conference 2016
LREC 2016main

Semantic Annotation of the ACL Anthology Corpus for the Automatic Analysis of Scientific Literature

Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016)

DOI:10.63317/2ktwfktn74ku

Abstract

This paper describes the process of creating a corpus annotated for concepts and semantic relations in the scientific domain. A part of the ACL Anthology Corpus was selected for annotation, but the annotation process itself is not specific to the computational linguistics domain and could be applied to any scientific corpora. Concepts were identified and annotated fully automatically, based on a combination of terminology extraction and available ontological resources. A typology of semantic relations between concepts is also proposed. This typology, consisting of 18 domain-specific and 3 generic relations, is the result of a corpus-based investigation of the text sequences occurring between concepts in sentences. A sample of 500 abstracts from the corpus is currently being manually annotated with these semantic relations. Only explicit relations are taken into account, so that the data could serve to train or evaluate pattern-based semantic relation classification systems.

Details

Paper ID
lrec2016-main-586
Pages
pp. 3694-3701
BibKey
gabor-etal-2016-semantic
Editor
N/A
Publisher
European Language Resources Association (ELRA)
ISSN
2522-2686
ISBN
978-2-9517408-9-1
Conference
Tenth International Conference on Language Resources and Evaluation
Location
Portorož, Slovenia
Date
23 May 2016 28 May 2016

Authors

  • KG

    Kata Gábor

  • HZ

    Haïfa Zargayouna

  • DB

    Davide Buscaldi

  • IT

    Isabelle Tellier

  • TC

    Thierry Charnois

Links