Automatic Term Extraction (ATE) aims to identify domain-specific lexical units, including single- and multi-word expressions, that encode the conceptual structure of specialized fields. This paper presents our submission to Subtask A (Term Extraction) of the ATE-IT shared task [1] at EVALITA 2026 [2], the first large-scale evaluation campaign dedicated to Italian ATE, focused on institutional texts in the municipal waste-management domain. To cope with sparse terminology, rich Italian morphology, and complex administrative multi-word terms, we adopt a hybrid pipeline that prioritizes recall and then progressively filters candidates. We first generate a broad pool of candidate terms using dependency-based heuristics with spaCy and zero-shot term identification via Google Gemini. We then apply normalization and task-specific constraints and refine predictions with a Random Forest classifier trained on a custom feature set capturing morphological patterns, domain-keyword density, and semantic similarity signals. We report official results on the test set and provide an analysis of typical errors, highlighting the benefits and limitations of combining linguistic heuristics, generative models, and supervised learning for Italian institutional automatic term extraction.
VVTE at ATE-IT: From Candidates to Terms: Hybrid Italian ATE with Dependency Heuristics, Gemini, and Random Forest Filtering
Di Nunzio G. M.
2026
Abstract
Automatic Term Extraction (ATE) aims to identify domain-specific lexical units, including single- and multi-word expressions, that encode the conceptual structure of specialized fields. This paper presents our submission to Subtask A (Term Extraction) of the ATE-IT shared task [1] at EVALITA 2026 [2], the first large-scale evaluation campaign dedicated to Italian ATE, focused on institutional texts in the municipal waste-management domain. To cope with sparse terminology, rich Italian morphology, and complex administrative multi-word terms, we adopt a hybrid pipeline that prioritizes recall and then progressively filters candidates. We first generate a broad pool of candidate terms using dependency-based heuristics with spaCy and zero-shot term identification via Google Gemini. We then apply normalization and task-specific constraints and refine predictions with a Random Forest classifier trained on a custom feature set capturing morphological patterns, domain-keyword density, and semantic similarity signals. We report official results on the test set and provide an analysis of typical errors, highlighting the benefits and limitations of combining linguistic heuristics, generative models, and supervised learning for Italian institutional automatic term extraction.Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.




