HomeLREC 2026WorkshopsPARLACLARINlrec2026-ws-parlaclarin-07
Back to PARLACLARIN 2026
LREC 2026workshop

Transcription and Recognition of Italian Parliamentary Speeches Using Vision-Language Models

Proceedings of the ParlaCLARIN V Workshop on Interoperability, Multilinguality, and Multimodality in Parliamentary Corpora

DOI:10.63317/587myq3zu4y2

Abstract

Parliamentary proceedings represent a rich yet challenging resource for computational analysis, particularly when preserved only as scanned historical documents. Existing efforts to digitise Italian parliamentary speeches have relied on traditional Optical Character Recognition pipelines, resulting in transcription errors and limited semantic annotation. In this paper, we propose a pipeline based on Vision-Language Models for the automatic transcription, semantic segmentation, and entity linking of Italian parliamentary speeches. The pipeline employs a specialised OCR model to extract text while preserving reading order, followed by a large-scale Vision-Language Model that performs transcription refinement, element classification, and speaker identification by jointly reasoning over visual layout and textual content. Extracted speakers are then linked to the Chamber of Deputies knowledge base through SPARQL queries and a multi-strategy fuzzy matching procedure. Evaluation against an established benchmark demonstrates substantial improvements both in transcription quality and speaker tagging.

Details

Paper ID
lrec2026-ws-parlaclarin-07
Pages
pp. 56-64
BibKey
curini-etal-2026-transcription
Editors
Maria Eskevich, Vincent Vandeghinste, David Bodron
Publisher
European Language Resources Association (ELRA)
ISSN
N/A
ISBN
N/A
Workshop
Proceedings of the ParlaCLARIN V Workshop on Interoperability, Multilinguality, and Multimodality in Parliamentary Corpora
Location
Palma, Mallorca, Spain
Date
11 - 16 May 2026

Authors

  • LC

    Luigi Curini

  • AF

    Alfio Ferrara

  • GP

    Giovanni Pagano

  • SP

    Sergio Picascia

Links