Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge

Young-Jun Lee; Dokyong Lee; Junyoung Youn; Kyeong-Jin Oh; Byungsoo Ko; Jonghwan Hyeon; Ho-Jin Choi

doi:10.18653/v1/2024.findings-emnlp.708

Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge

Young-Jun Lee, Dokyong Lee, Junyoung Youn, Kyeong-Jin Oh, Byungsoo Ko, Jonghwan Hyeon, Ho-Jin Choi

Abstract

Humans share a wide variety of images related to their personal experiences within conversations via instant messaging tools. However, existing works focus on (1) image-sharing behavior in singular sessions, leading to limited long-term social interaction, and (2) a lack of personalized image-sharing behavior. In this work, we introduce , a large-scale long-term multi-modal dialogue dataset that covers a wide range of social personas in a multi-modality format, time intervals, and images. To construct automatically, we propose a novel multi-modal contextualization framework, , that generates long-term multi-modal dialogue distilled from ChatGPT and our proposed image aligner. Using our , we train a multi-modal conversation model, 7B, which demonstrates impressive visual imagination ability. Furthermore, we demonstrate the effectiveness of our dataset in human evaluation. The code, dataset, and model will be publicly released after publication.

Anthology ID:: 2024.findings-emnlp.708
Volume:: Findings of the Association for Computational Linguistics: EMNLP 2024
Month:: November
Year:: 2024
Address:: Miami, Florida, USA
Editors:: Yaser Al-Onaizan, Mohit Bansal, Yun-Nung Chen
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 12137–12162
Language:
URL:: https://aclanthology.org/2024.findings-emnlp.708/
DOI:: 10.18653/v1/2024.findings-emnlp.708
Bibkey:
Cite (ACL):: Young-Jun Lee, Dokyong Lee, Junyoung Youn, Kyeong-Jin Oh, Byungsoo Ko, Jonghwan Hyeon, and Ho-Jin Choi. 2024. Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 12137–12162, Miami, Florida, USA. Association for Computational Linguistics.
Cite (Informal):: Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge (Lee et al., Findings 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.findings-emnlp.708.pdf

PDF Cite Search Fix data