Learning Zero-Shot Multifaceted Visually Grounded Word Embeddings via Multi-Task Training

Hassan Shahmohammadi; Hendrik P. A. Lensch; Harald Baayen

doi:10.18653/v1/2021.conll-1.12

Learning Zero-Shot Multifaceted Visually Grounded Word Embeddings via Multi-Task Training

Hassan Shahmohammadi, Hendrik P. A. Lensch, R. Harald Baayen

Abstract

Language grounding aims at linking the symbolic representation of language (e.g., words) into the rich perceptual knowledge of the outside world. The general approach is to embed both textual and visual information into a common space -the grounded space- confined by an explicit relationship. We argue that since concrete and abstract words are processed differently in the brain, such approaches sacrifice the abstract knowledge obtained from textual statistics in the process of acquiring perceptual information. The focus of this paper is to solve this issue by implicitly grounding the word embeddings. Rather than learning two mappings into a joint space, our approach integrates modalities by implicit alignment. This is achieved by learning a reversible mapping between the textual and the grounded space by means of multi-task training. Intrinsic and extrinsic evaluations show that our way of visual grounding is highly beneficial for both abstract and concrete words. Our embeddings are correlated with human judgments and outperform previous works using pretrained word embeddings on a wide range of benchmarks. Our grounded embeddings are publicly available here.

Anthology ID:: 2021.conll-1.12
Volume:: Proceedings of the 25th Conference on Computational Natural Language Learning
Month:: November
Year:: 2021
Address:: Online
Editors:: Arianna Bisazza, Omri Abend
Venue:: CoNLL
SIG:: SIGNLL
Publisher:: Association for Computational Linguistics
Note:
Pages:: 158–170
Language:
URL:: https://aclanthology.org/2021.conll-1.12
DOI:: 10.18653/v1/2021.conll-1.12
Bibkey:
Cite (ACL):: Hassan Shahmohammadi, Hendrik P. A. Lensch, and R. Harald Baayen. 2021. Learning Zero-Shot Multifaceted Visually Grounded Word Embeddings via Multi-Task Training. In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 158–170, Online. Association for Computational Linguistics.
Cite (Informal):: Learning Zero-Shot Multifaceted Visually Grounded Word Embeddings via Multi-Task Training (Shahmohammadi et al., CoNLL 2021)
Copy Citation:
PDF:: https://aclanthology.org/2021.conll-1.12.pdf
Video:: https://aclanthology.org/2021.conll-1.12.mp4
Code: Hazel1994/Visually_Grounded_Word_Embeddings
Data: MS COCO

PDF Cite Search Code Video