Automatic Term Extraction (ATE) is a core task in Natural Language Processing, aiming to identify domain-specific terminology that can support downstream applications such as information retrieval, ontology construction, and domain adaptation. This paper presents a sequence labeling approach to the ATE-IT shared task at EVALITA 2026, focusing on the extraction of terms from Italian institutional texts in the waste management domain. The task is formulated as a BIO tagging problem and addressed using a linear-chain Conditional Random Field (CRF) model. The proposed system integrates symbolic linguistic features obtained from spaCy, contextual semantic representations derived from a pretrained Italian BERT model, and weak domain knowledge through an automatically constructed glossary extracted from the training data. BERT embeddings are used as fixed contextual features and combined with token-level linguistic information within the CRF framework to enforce coherent span-level predictions. Experimental results on both development and test sets show that the proposed approach consistently improves Type F1-score over the baseline, with particularly notable gains in precision, indicating more accurate reconstruction of complete term boundaries. While recall remains lower than the baseline, the analysis highlights the effectiveness of combining contextual embeddings with structured decoding for Italian Automatic Term Extraction and outlines directions for improving generalization and recall in future work.

JTTE at ATE-IT: A CRF Model with Contextual Embeddings

Di Nunzio G. M.
2026

Abstract

Automatic Term Extraction (ATE) is a core task in Natural Language Processing, aiming to identify domain-specific terminology that can support downstream applications such as information retrieval, ontology construction, and domain adaptation. This paper presents a sequence labeling approach to the ATE-IT shared task at EVALITA 2026, focusing on the extraction of terms from Italian institutional texts in the waste management domain. The task is formulated as a BIO tagging problem and addressed using a linear-chain Conditional Random Field (CRF) model. The proposed system integrates symbolic linguistic features obtained from spaCy, contextual semantic representations derived from a pretrained Italian BERT model, and weak domain knowledge through an automatically constructed glossary extracted from the training data. BERT embeddings are used as fixed contextual features and combined with token-level linguistic information within the CRF framework to enforce coherent span-level predictions. Experimental results on both development and test sets show that the proposed approach consistently improves Type F1-score over the baseline, with particularly notable gains in precision, indicating more accurate reconstruction of complete term boundaries. While recall remains lower than the baseline, the analysis highlights the effectiveness of combining contextual embeddings with structured decoding for Italian Automatic Term Extraction and outlines directions for improving generalization and recall in future work.
2026
CEUR Workshop Proceedings
9th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop, EVALITA 2026
File in questo prodotto:
Non ci sono file associati a questo prodotto.
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11577/3609611
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact