%0 Conference Proceedings %T “Caption” as a Coherence Relation: Evidence and Implications %A Alikhani, Malihe %A Stone, Matthew %Y Bernardi, Raffaella %Y Fernandez, Raquel %Y Gella, Spandana %Y Kafle, Kushal %Y Kanan, Christopher %Y Lee, Stefan %Y Nabi, Moin %S Proceedings of the Second Workshop on Shortcomings in Vision and Language %D 2019 %8 June %I Association for Computational Linguistics %C Minneapolis, Minnesota %F alikhani-stone-2019-caption %X We study verbs in image–text corpora, contrasting caption corpora, where texts are explicitly written to characterize image content, with depiction corpora, where texts and images may stand in more general relations. Captions show a distinctively limited distribution of verbs, with strong preferences for specific tense, aspect, lexical aspect, and semantic field. These limitations, which appear in data elicited by a range of methods, restrict the utility of caption corpora to inform image retrieval, multimodal document generation, and perceptually-grounded semantic models. We suggest that these limitations reflect the discourse constraints in play when subjects write texts to accompany imagery, so we argue that future development of image–text corpora should work to increase the diversity of event descriptions, while looking explicitly at the different ways text and imagery can be coherently related. %R 10.18653/v1/W19-1806 %U https://aclanthology.org/W19-1806 %U https://doi.org/10.18653/v1/W19-1806 %P 58-67