VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP

Tu Tran Do, Nhat Ngoc Nguyen, Tung Khanh Tran, Hoang D. Nguyen, Tu Minh Phuong, Long Hoang Dang


Abstract
We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese. VIVID comprises 1,636 idioms and proverbs annotated with five complexity traits (literal expressions, pragmatic nuances, Sino-Vietnamese terms, uncommon vocabulary, folk knowledge) and seven semantic themes. We establish an evaluation framework combining generative and discriminative tasks, proposing an LLM-as-a-Judge approach with aspect-based prompting validated against human judgment (Cohen’s κ = 0.792). Evaluating eight state-of-the-art models reveals critical gaps: Vietnamese-specialized models drastically underperform multilingual systems (VinaLLaMA-7B: 0.13 vs. GPT-4o: 2.46), and even top models achieve less than 50% of maximum scores. Notably, few-shot prompting does not universally improve performance, with GPT-4o exhibiting degradation due to stylistic overfitting. Our analysis exposes systematic failures including literal over-interpretation, lexical gaps, and pragmatic flattening, demonstrating that current models lack cultural competence for nuanced figurative interpretation. VIVID provides an essential tool for advancing figurative language understanding in culturally rich contexts.
Anthology ID:
2026.lrec-1.422
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
5414–5430
Language:
External URL:
https://lrec.elra.info/lrec2026-main-422
DOI:
10.63317/3b4ag6rpijwb
Bibkey:
Cite (ACL):
Tu Tran Do, Nhat Ngoc Nguyen, Tung Khanh Tran, Hoang D. Nguyen, Tu Minh Phuong, and Long Hoang Dang. 2026. VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 5414–5430, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP (Do et al., LREC 2026)
Copy Citation: