Marco Forlano


2024

pdf bib
Towards the WhAP Corpus: A Resource for the Study of Italian on WhatsApp
Ilaria Fiorentini | Marco Forlano | Nicholas Nese
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)

Over the past two decades, the rise of new technologies and social networks has significantly shaped written language, imbuing it with characteristics akin to the spoken language. This study reports on the ongoing initiative to build the WhAP corpus, a resource featuring WhatsApp conversations in Italian, encompassing both written and spoken messages and totaling at present more than 400.000 tokens, 89 conversations, and 194 participants from diverse age groups and geographical regions of Italy. More specifically, this paper focuses on the practical steps involved in the construction of the resource. Once publicly accessible, the WhAP Corpus will enable in-depth linguistic research on the language used on WhatsApp, which shows unique features such as the blending of written and spoken elements.

2023

pdf bib
LSDT: a Dependency Treebank of Lombard Sinti
Marco Forlano | Luca Brigada Villa
Proceedings of the Sixth Workshop on the Use of Computational Methods in the Study of Endangered Languages