Chiara Fazzone


2024

pdf bib
SimilEx: The First Italian Dataset for Sentence Similarity with Natural Language Explanations
Chiara Alzetta | Felice Dell’orletta | Chiara Fazzone | Giulia Venturi
Proceedings of the 10th Italian Conference on Computational Linguistics (CLiC-it 2024)

Large language models (LLMs) demonstrate great performance in natural language processing and understanding tasks. However, much work remains to enhance their interpretability. Annotated datasets with explanations could be key to addressing this issue, as they enable the development of models that provide human-like explanations for their decisions. In this paper, we introduce the SimilEx dataset, the first Italian dataset reporting human evaluations of similarity between pairs of sentences. For a subset of these pairs, the annotators also provided explanations in natural language for the scores assigned. The SimilEx dataset is valuable for exploring the variability in similarity perception between sentences and for training LLMs to offer human-like explanations for their predictions.