Calogero Jerik Scozzaro
2026
ProverbIT - Easy to complete, hard to choose: A CALAMITA Challenge
Enrico Mensa | Lorenzo Zane | Calogero Jerik Scozzaro | Matteo Delsanto | Tommaso Milani | Daniele P. Radicioni
Proceedings of the Ninth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA 2026)
Enrico Mensa | Lorenzo Zane | Calogero Jerik Scozzaro | Matteo Delsanto | Tommaso Milani | Daniele P. Radicioni
Proceedings of the Ninth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA 2026)
Kenji-Endo: a BabyLM @EVALITA
Calogero Jerik Scozzaro | Matteo Rinaldi | Gianluca Mittone | Marco Antonio Stranisci
Proceedings of the Ninth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA 2026)
Calogero Jerik Scozzaro | Matteo Rinaldi | Gianluca Mittone | Marco Antonio Stranisci
Proceedings of the Ninth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA 2026)
2025
Beyond the Average Reader: the Reader Embedding Approach
Calogero Jerik Scozzaro | Matteo Delsanto | Daniele P. Radicioni
Findings of the Association for Computational Linguistics: ACL 2025
Calogero Jerik Scozzaro | Matteo Delsanto | Daniele P. Radicioni
Findings of the Association for Computational Linguistics: ACL 2025
Focus of this work is the prediction of reading times as the task is customarily dealt with in literature: that is, by collecting eye-tracking data that are averaged and employed to train learning models. We start by observing that systems trained on average values are ill-suited for the prediction of the reading times for specific subjects, as they fail to account for individual variability and accurately analyze the reading gestures of specific reader groups, or to target specific user needs. To overcome such limitation, that is to predict the reading times for a specific subject, we propose a novel approach based on creating an embedding to compactly describe her/his fixations. Embeddings are used to individuate readers that share same or similar reading behavior from a reference corpus. Models are then trained on values averaged over this subset of similar readers. Experimental results indicate that the proposed approach consistently outperforms its corresponding variants, in which predictions of reading times for specific readers are based on data from all subjects rather than from the most similar ones.