Plot Twist: Multimodal Models Don’t Comprehend Simple Chart Details

Yasaman Razeghi; Ishita Dasgupta; Fangyu Liu; Vinay Ramasesh; Sameer Singh

Plot Twist: Multimodal Models Don’t Comprehend Simple Chart Details

Yasaman Razeghi, Ishita Dasgupta, Fangyu Liu, Vinay Ramasesh, Sameer Singh

Abstract

Recent advances in multimodal models show remarkable performance in real-world benchmarks for chart and figure understanding like ChartQA that involve interpreting trends, comparing data points, and extracting insights from visuals.In this paper, we investigate the extent to which these models truly comprehend the underlying information in charts by posing direct, elementary questions about simple features such as axes ranges and values to examine their fundamental visual understanding abilities in the context of charts.Our questions are applied to two sets of figures: synthetic and real-world.The empirical evaluation of 5 popular multimodal models on our dataset reveals shortfalls in understanding charts and figures, contrary to what their performance on complex benchmarks might suggest.For instance, Gemini Pro Vision only achieves 57.9% accuracy on our elementary set of questions on real-world plots, while other popular multimodal models showed similar or less performance.This work highlights an important limitation of current multimodal models, and cautions against overly optimistic interpretations of their abilities based on results of canonical evaluations.

Anthology ID:: 2024.findings-emnlp.342
Volume:: Findings of the Association for Computational Linguistics: EMNLP 2024
Month:: November
Year:: 2024
Address:: Miami, Florida, USA
Editors:: Yaser Al-Onaizan, Mohit Bansal, Yun-Nung Chen
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 5922–5937
Language:
URL:: https://aclanthology.org/2024.findings-emnlp.342
DOI:
Bibkey:
Cite (ACL):: Yasaman Razeghi, Ishita Dasgupta, Fangyu Liu, Vinay Ramasesh, and Sameer Singh. 2024. Plot Twist: Multimodal Models Don’t Comprehend Simple Chart Details. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 5922–5937, Miami, Florida, USA. Association for Computational Linguistics.
Cite (Informal):: Plot Twist: Multimodal Models Don’t Comprehend Simple Chart Details (Razeghi et al., Findings 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.findings-emnlp.342.pdf

PDF Cite Search