Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models

Srishti Yadav; Zhi Zhang; Daniel Hershcovich; Ekaterina Shutova

doi:10.18653/v1/2025.findings-naacl.422

Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models

Srishti Yadav, Zhi Zhang, Daniel Hershcovich, Ekaterina Shutova

Abstract

Investigating value alignment in Large Language Models (LLMs) based on cultural context has become a critical area of research. However, similar biases have not been extensively explored in large vision-language models (VLMs). As the scale of multimodal models continues to grow, it becomes increasingly important to assess whether images can serve as reliable proxies for culture and how these values are embedded through the integration of both visual and textual data. In this paper, we conduct a thorough evaluation of multimodal model at different scales, focusing on their alignment with cultural values. Our findings reveal that, much like LLMs, VLMs exhibit sensitivity to cultural values, but their performance in aligning with these values is highly context-dependent. While VLMs show potential in improving value understanding through the use of images, this alignment varies significantly across contexts highlighting the complexities and underexplored challenges in the alignment of multimodal models.

Anthology ID:: 2025.findings-naacl.422
Volume:: Findings of the Association for Computational Linguistics: NAACL 2025
Month:: April
Year:: 2025
Address:: Albuquerque, New Mexico
Editors:: Luis Chiruzzo, Alan Ritter, Lu Wang
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 7607–7623
Language:
URL:: https://aclanthology.org/2025.findings-naacl.422/
DOI:: 10.18653/v1/2025.findings-naacl.422
Bibkey:
Cite (ACL):: Srishti Yadav, Zhi Zhang, Daniel Hershcovich, and Ekaterina Shutova. 2025. Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 7607–7623, Albuquerque, New Mexico. Association for Computational Linguistics.
Cite (Informal):: Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models (Yadav et al., Findings 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.findings-naacl.422.pdf

PDF Cite Search Fix data