BAN-Cap: A Multi-Purpose English-Bangla Image Descriptions Dataset

Mohammad Faiyaz Khan, S.M. Sadiq-Ur-Rahman Shifath, Md Saiful Islam


Abstract
As computers have become efficient at understanding visual information and transforming it into a written representation, research interest in tasks like automatic image captioning has seen a significant leap over the last few years. While most of the research attention is given to the English language in a monolingual setting, resource-constrained languages like Bangla remain out of focus, predominantly due to a lack of standard datasets. Addressing this issue, we present a new dataset BAN-Cap following the widely used Flickr8k dataset, where we collect Bangla captions of the images provided by qualified annotators. Our dataset represents a wider variety of image caption styles annotated by trained people from different backgrounds. We present a quantitative and qualitative analysis of the dataset and the baseline evaluation of the recent models in Bangla image captioning. We investigate the effect of text augmentation and demonstrate that an adaptive attention-based model combined with text augmentation using Contextualized Word Replacement (CWR) outperforms all state-of-the-art models for Bangla image captioning. We also present this dataset’s multipurpose nature, especially on machine translation for Bangla-English and English-Bangla. This dataset and all the models will be useful for further research.
Anthology ID:
2022.lrec-1.740
Volume:
Proceedings of the Thirteenth Language Resources and Evaluation Conference
Month:
June
Year:
2022
Address:
Marseille, France
Editors:
Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association
Note:
Pages:
6855–6865
Language:
URL:
https://aclanthology.org/2022.lrec-1.740
DOI:
Bibkey:
Cite (ACL):
Mohammad Faiyaz Khan, S.M. Sadiq-Ur-Rahman Shifath, and Md Saiful Islam. 2022. BAN-Cap: A Multi-Purpose English-Bangla Image Descriptions Dataset. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 6855–6865, Marseille, France. European Language Resources Association.
Cite (Informal):
BAN-Cap: A Multi-Purpose English-Bangla Image Descriptions Dataset (Khan et al., LREC 2022)
Copy Citation:
PDF:
https://aclanthology.org/2022.lrec-1.740.pdf
Code
 faiyazkhan11/ban-cap