Automatic Term Extraction (ATE) is a core task in Natural Language Processing, aiming to identify domain-specific terminology that can support downstream applications such as information retrieval, ontology construction, and domain adaptation. This paper presents a sequence labeling approach to the ATE-IT shared task at EVALITA 2026, focusing on the extraction of terms from Italian institutional texts in the waste management domain. The task is formulated as a BIO tagging problem and addressed using a linear-chain Conditional Random Field (CRF) model. The proposed system integrates symbolic linguistic features obtained from spaCy, contextual semantic representations derived from a pretrained Italian BERT model, and weak domain knowledge through an automatically constructed glossary extracted from the training data. BERT embeddings are used as fixed contextual features and combined with token-level linguistic information within the CRF framework to enforce coherent span-level predictions. Experimental results on both development and test sets show that the proposed approach consistently improves Type F1-score over the baseline, with particularly notable gains in precision, indicating more accurate reconstruction of complete term boundaries. While recall remains lower than the baseline, the analysis highlights the effectiveness of combining contextual embeddings with structured decoding for Italian Automatic Term Extraction and outlines directions for improving generalization and recall in future work.
JTTE at ATE-IT: A CRF Model with Contextual Embeddings
Di Nunzio G. M.
2026
Abstract
Automatic Term Extraction (ATE) is a core task in Natural Language Processing, aiming to identify domain-specific terminology that can support downstream applications such as information retrieval, ontology construction, and domain adaptation. This paper presents a sequence labeling approach to the ATE-IT shared task at EVALITA 2026, focusing on the extraction of terms from Italian institutional texts in the waste management domain. The task is formulated as a BIO tagging problem and addressed using a linear-chain Conditional Random Field (CRF) model. The proposed system integrates symbolic linguistic features obtained from spaCy, contextual semantic representations derived from a pretrained Italian BERT model, and weak domain knowledge through an automatically constructed glossary extracted from the training data. BERT embeddings are used as fixed contextual features and combined with token-level linguistic information within the CRF framework to enforce coherent span-level predictions. Experimental results on both development and test sets show that the proposed approach consistently improves Type F1-score over the baseline, with particularly notable gains in precision, indicating more accurate reconstruction of complete term boundaries. While recall remains lower than the baseline, the analysis highlights the effectiveness of combining contextual embeddings with structured decoding for Italian Automatic Term Extraction and outlines directions for improving generalization and recall in future work.Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.




