Scene Graph Enhanced Pseudo-Labeling for Referring Expression Comprehension

Cantao Wu, Yi Cai, Liuwu Li, Jiexin Wang


Abstract
Referring Expression Comprehension (ReC) is a task that involves localizing objects in images based on natural language expressions. Most ReC methods typically approach the task as a supervised learning problem. However, the need for costly annotations, such as clear image-text pairs or region-text pairs, hinders the scalability of existing approaches. In this work, we propose a novel scene graph-based framework that automatically generates high-quality pseudo region-query pairs. Our method harnesses scene graphs to capture the relationships between objects in images and generate expressions enriched with relation information. To ensure accurate mapping between visual regions and text, we introduce an external module that employs a calibration algorithm to filter out ambiguous queries. Additionally, we employ a rewriter module to enhance the diversity of our generated pseudo queries through rewriting. Extensive experiments demonstrate that our method outperforms previous pseudo-labeling methods by about 10%, 12%, and 11% on RefCOCO, RefCOCO+, and RefCOCOg, respectively. Furthermore, it surpasses the state-of-the-art unsupervised approach by more than 15% on the RefCOCO dataset.
Anthology ID:
2023.findings-emnlp.802
Volume:
Findings of the Association for Computational Linguistics: EMNLP 2023
Month:
December
Year:
2023
Address:
Singapore
Editors:
Houda Bouamor, Juan Pino, Kalika Bali
Venue:
Findings
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
11978–11990
Language:
URL:
https://aclanthology.org/2023.findings-emnlp.802
DOI:
10.18653/v1/2023.findings-emnlp.802
Bibkey:
Cite (ACL):
Cantao Wu, Yi Cai, Liuwu Li, and Jiexin Wang. 2023. Scene Graph Enhanced Pseudo-Labeling for Referring Expression Comprehension. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 11978–11990, Singapore. Association for Computational Linguistics.
Cite (Informal):
Scene Graph Enhanced Pseudo-Labeling for Referring Expression Comprehension (Wu et al., Findings 2023)
Copy Citation:
PDF:
https://aclanthology.org/2023.findings-emnlp.802.pdf