Back to Main Conference 2024
LREC-COLING 2024main

Lemmatisation of Medieval Greek: Against the Limits of Transformer’s Capabilities?

Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)

DOI:10.63317/35evzb2ctsc3

Abstract

This paper presents preliminary experiments for the lemmatisation of unedited, Byzantine Greek epigrams. This type of Greek is quite different from its classical ancestor, mostly because of its orthographic inconsistencies. Existing lemmatisation algorithms display an accuracy drop of around 30pp when tested on these Byzantine book epigrams. We conducted seven different lemmatisation experiments, which were either transformer-based or based on neural edit-trees. The best performing lemmatiser was a hybrid method combining transformer-based embeddings with a dictionary look-up. We compare our results with existing lemmatisers, and provide a detailed error analysis revealing why unedited, Byzantine Greek is so challenging for lemmatisation.

Details

Paper ID
lrec2024-main-0899
Pages
pp. 10293-10302
BibKey
swaelens-etal-2024-lemmatisation
Editor
N/A
Publisher
European Language Resources Association (ELRA) and ICCL
ISSN
2522-2686
ISBN
979-10-95546-34-4
Conference
Joint International Conference on Computational Linguistics, Language Resources and Evaluation
Location
Turin, Italy
Date
20 May 2024 25 May 2024

Authors

  • CS

    Colin Swaelens

  • PS

    Pranaydeep Singh

  • Id

    Ilse de Vos

  • EL

    Els Lefever

Links