DREAM: A Multicultural Multimodal Dataset Linking Dialogues and Realistic Image Sequences

Juan Mallo, Marcos Estecha-Garitagoitia, Ricardo Cordoba, Luis Fernando D’Haro


Abstract
An ongoing challenge in multimodal language research is creating and interpreting dialogues that preserve visual and cultural consistency across turns. We introduce DREAM (Dialogue to REAlistic Multicultural Image Sequences), a multicultural multimodal resource that ties dialogues grounded in explicit persona profiles to photorealistic, storyboard-like image sequences. Each of the 1,000 dialogues includes two rich persona profiles (structured traits plus descriptive language), two matching photorealistic portraits, and a collection of scene-level images depicting key dialogue moments. The pipeline integrates profile augmentation, culturally-sensitive prompt engineering, and turn selection to craft cohesive visual narratives, promoting character consistency across images. This is accomplished through a controlled generation process employing large language and image models. Beyond dialogue grounding, DREAM supports appearance-based demographic perception and culture-aware rendering: models can be evaluated on their ability to (i) perceive age, gender presentation, and broad ethnicity appearance clusters from profile portraits, and (ii) maintain these characteristics in dialogue scenes. We provide a unified JSON format integrating profiles, dialogue text, and visual turns, facilitating research on visually anchored dialogue understanding, consistency, and generation. A dual evaluation protocol combines human judgments (realism, coherence, consistency, and demographic perception) with automated portrait analysis via GPT-5. Ethical considerations, limitations, and recommended applications are discussed.
Anthology ID:
2026.lrec-1.728
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
9266–9281
Language:
External URL:
https://lrec.elra.info/lrec2026-main-728
DOI:
10.63317/2v7b4xhs2d5g
Bibkey:
Cite (ACL):
Juan Mallo, Marcos Estecha-Garitagoitia, Ricardo Cordoba, and Luis Fernando D’Haro. 2026. DREAM: A Multicultural Multimodal Dataset Linking Dialogues and Realistic Image Sequences. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 9266–9281, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
DREAM: A Multicultural Multimodal Dataset Linking Dialogues and Realistic Image Sequences (Mallo et al., LREC 2026)
Copy Citation: