Filippo Merlo
Author directory2026
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
Filippo Merlo | Ece Takmaz | Wenkai Chen | Albert Gatt
Transactions of the Association for Computational Linguistics, Volume 14
Filippo Merlo | Ece Takmaz | Wenkai Chen | Albert Gatt
Transactions of the Association for Computational Linguistics, Volume 14
To what degree and under what conditions do VLMs rely on scene context when generating references to objects? To address this question, we introduce the Common Objects Out-of-Context (COOCo) dataset and conduct experiments on several VLMs under different degrees of scene–object congruency and noise. We find that models leverage scene context adaptively, depending on scene-object semantic relatedness and noise level. Based on these consistent trends across models, we turn to the question of how VLM attention patterns change as a function of target-scene semantic fit, and to what degree these patterns are predictive of categorisation accuracy. We find that successful object categorisation is associated with increased mid-layer attention to the target. We also find a non-monotonic dependency on semantic fit, with attention dropping at moderate fit and increasing for both low and high fit. These results suggest that VLMs dynamically balance local and contextual information for reference generation. Dataset and code are available here: https://github.com/cs-nlp-uu/scenereg.
2023
ChatGPT’s Information Seeking Strategy: Insights from the 20-Questions Game
Leonardo Bertolazzi | Davide Mazzaccara | Filippo Merlo | Raffaella Bernardi
Proceedings of the 16th International Natural Language Generation Conference
Leonardo Bertolazzi | Davide Mazzaccara | Filippo Merlo | Raffaella Bernardi
Proceedings of the 16th International Natural Language Generation Conference
Large Language Models, and ChatGPT in particular, have recently grabbed the attention of the community and the media. Having reached high language proficiency, attention has been shifting toward its reasoning capabilities. In this paper, our main aim is to evaluate ChatGPT’s question generation in a task where language production should be driven by an implicit reasoning process. To this end, we employ the 20-Questions game, traditionally used within the Cognitive Science community to inspect the information seeking-strategy’s development. This task requires a series of interconnected skills: asking informative questions, stepwise updating the hypothesis space, and stopping asking questions when enough information has been collected. We build hierarchical hypothesis spaces, exploiting feature norms collected from humans vs. ChatGPT itself, and we inspect the efficiency and informativeness of ChatGPT’s strategy. Our results show that ChatGPT’s performance gets closer to an optimal agent only when prompted to explicitly list the updated space stepwise.