Franca Debole


pdf bib
Italian Word Embeddings for the Medical Domain
Franco Alberto Cardillo | Franca Debole
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)

Neural word embeddings have proven valuable in the development of medical applications. However, for the Italian language, there are no publicly available corpora, embeddings, or evaluation resources tailored to this domain. In this paper, we introduce an Italian corpus for the medical domain, that includes texts from Wikipedia, medical journals, drug leaflets, and specialized websites. Using this corpus, we generate neural word embeddings from scratch. These embeddings are then evaluated using standard evaluation resources, that we translated into Italian exploiting the concept graph in the UMLS Metathesaurus. Despite the relatively small size of the corpus, our experimental results indicate that the new embeddings correlate well with human judgments regarding the similarity and the relatedness of medical concepts. Moreover, these medical-specific embeddings outperform a baseline model trained on the full Wikipedia corpus, which includes the medical pages we used. We believe that our embeddings and the newly introduced textual resources will foster further advancements in the field of Italian medical Natural Language Processing.


pdf bib
Multilingual Search for Cultural Heritage Archives via Combining Multiple Translation Resources
Gareth J. F. Jones | Ying Zhang | Eamonn Newman | Fabio Fantino | Franca Debole
Proceedings of the Workshop on Language Technology for Cultural Heritage Data (LaTeCH 2007).


pdf bib
An Analysis of the Relative Difficulty of Reuters-21578 Subsets
Franca Debole | Fabrizio Sebastiani
Proceedings of the Fourth International Conference on Language Resources and Evaluation (LREC’04)