This paper presents our participation in Subtask A of the ATE-IT shared task at EVALITA 2026, which targets sentence-level Automatic Terminology Extraction from Italian administrative documents in the waste-management domain. The dataset includes a balanced collection of administrative acts (e.g., municipal regulations, service charters, tenders) and informative texts (e.g., public notices), combining formal institutional language with less rigid informational prose. These heterogeneous documents are characterized by syntactically complex sentences and dense specialized terminology. Given an input sentence, systems must identify all domain-relevant single- and multiword terms and output them in a predefined JSON format, requiring generalization beyond a fixed vocabulary. We investigate and compare two strategies: (i) a supervised neural approach based on an Italian BERT model fine-tuned for token-level term span detection, and (ii) a zero-shot approach using large language models to extract terminology without task-specific training. We report official results and provide an analysis of typical errors to highlight the relative strengths and limitations of supervised versus zero-shot methods for terminology extraction in complex institutional texts.

DSBKTE at ATE-IT: From Token Classification to Zero-Shot Generation: Two Approaches to Italian ATE at EVALITA 2026

Di Nunzio G. M.
2026

Abstract

This paper presents our participation in Subtask A of the ATE-IT shared task at EVALITA 2026, which targets sentence-level Automatic Terminology Extraction from Italian administrative documents in the waste-management domain. The dataset includes a balanced collection of administrative acts (e.g., municipal regulations, service charters, tenders) and informative texts (e.g., public notices), combining formal institutional language with less rigid informational prose. These heterogeneous documents are characterized by syntactically complex sentences and dense specialized terminology. Given an input sentence, systems must identify all domain-relevant single- and multiword terms and output them in a predefined JSON format, requiring generalization beyond a fixed vocabulary. We investigate and compare two strategies: (i) a supervised neural approach based on an Italian BERT model fine-tuned for token-level term span detection, and (ii) a zero-shot approach using large language models to extract terminology without task-specific training. We report official results and provide an analysis of typical errors to highlight the relative strengths and limitations of supervised versus zero-shot methods for terminology extraction in complex institutional texts.
2026
CEUR Workshop Proceedings
9th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop, EVALITA 2026
File in questo prodotto:
Non ci sono file associati a questo prodotto.
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11577/3609610
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact