Automatic Term Extraction (ATE) aims to identify domain-specific lexical units, including single- and multi-word expressions, that encode the conceptual structure of specialized fields. This paper presents our submission to Subtask A (Term Extraction) of the ATE-IT shared task [1] at EVALITA 2026 [2], the first large-scale evaluation campaign dedicated to Italian ATE, focused on institutional texts in the municipal waste-management domain. To cope with sparse terminology, rich Italian morphology, and complex administrative multi-word terms, we adopt a hybrid pipeline that prioritizes recall and then progressively filters candidates. We first generate a broad pool of candidate terms using dependency-based heuristics with spaCy and zero-shot term identification via Google Gemini. We then apply normalization and task-specific constraints and refine predictions with a Random Forest classifier trained on a custom feature set capturing morphological patterns, domain-keyword density, and semantic similarity signals. We report official results on the test set and provide an analysis of typical errors, highlighting the benefits and limitations of combining linguistic heuristics, generative models, and supervised learning for Italian institutional automatic term extraction.

VVTE at ATE-IT: From Candidates to Terms: Hybrid Italian ATE with Dependency Heuristics, Gemini, and Random Forest Filtering

Di Nunzio G. M.
2026

Abstract

Automatic Term Extraction (ATE) aims to identify domain-specific lexical units, including single- and multi-word expressions, that encode the conceptual structure of specialized fields. This paper presents our submission to Subtask A (Term Extraction) of the ATE-IT shared task [1] at EVALITA 2026 [2], the first large-scale evaluation campaign dedicated to Italian ATE, focused on institutional texts in the municipal waste-management domain. To cope with sparse terminology, rich Italian morphology, and complex administrative multi-word terms, we adopt a hybrid pipeline that prioritizes recall and then progressively filters candidates. We first generate a broad pool of candidate terms using dependency-based heuristics with spaCy and zero-shot term identification via Google Gemini. We then apply normalization and task-specific constraints and refine predictions with a Random Forest classifier trained on a custom feature set capturing morphological patterns, domain-keyword density, and semantic similarity signals. We report official results on the test set and provide an analysis of typical errors, highlighting the benefits and limitations of combining linguistic heuristics, generative models, and supervised learning for Italian institutional automatic term extraction.
2026
CEUR Workshop Proceedings
9th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop, EVALITA 2026
File in questo prodotto:
Non ci sono file associati a questo prodotto.
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11577/3609605
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact