Back to Main Conference 2014
LREC 2014main

Priberam Compressive Summarization Corpus: A New Multi-Document Summarization Corpus for European Portuguese

Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC 2014)

DOI:10.63317/5mnnogd54oo2

Abstract

In this paper, we introduce the Priberam Compressive Summarization Corpus, a new multi-document summarization corpus for European Portuguese. The corpus follows the format of the summarization corpora for English in recent DUC and TAC conferences. It contains 80 manually chosen topics referring to events occurred between 2010 and 2013. Each topic contains 10 news stories from major Portuguese newspapers, radio and TV stations, along with two human generated summaries up to 100 words. Apart from the language, one important difference from the DUC/TAC setup is that the human summaries in our corpus are compressive: the annotators performed only sentence and word deletion operations, as opposed to generating summaries from scratch. We use this corpus to train and evaluate learning-based extractive and compressive summarization systems, providing an empirical comparison between these two approaches. The corpus is made freely available in order to facilitate research on automatic summarization.

Details

Paper ID
lrec2014-main-193
Pages
pp. 146-152
BibKey
almeida-etal-2014-priberam
Editor
N/A
Publisher
European Language Resources Association (ELRA)
ISSN
2522-2686
ISBN
978-2-9517408-8-4
Conference
Ninth International Conference on Language Resources and Evaluation
Location
Reykjavik, Iceland
Date
26 May 2014 31 May 2014

Authors

  • MA

    Miguel B. Almeida

  • MA

    Mariana S. C. Almeida

  • AM

    André F. T. Martins

  • HF

    Helena Figueira

  • PM

    Pedro Mendes

  • CP

    Cláudia Pinto

Links