No Gestures Left Behind: Learning Relationships between Spoken Language and Freeform Gestures

Chaitanya Ahuja, Dong Won Lee, Ryo Ishii, Louis-Philippe Morency


Abstract
We study relationships between spoken language and co-speech gestures in context of two key challenges. First, distributions of text and gestures are inherently skewed making it important to model the long tail. Second, gesture predictions are made at a subword level, making it important to learn relationships between language and acoustic cues. We introduce AISLe, which combines adversarial learning with importance sampling to strike a balance between precision and coverage. We propose the use of a multimodal multiscale attention block to perform subword alignment without the need of explicit alignment between language and acoustic cues. Finally, to empirically study the importance of language in this task, we extend the dataset proposed in Ahuja et al. (2020) with automatically extracted transcripts for audio signals. We substantiate the effectiveness of our approach through large-scale quantitative and user studies, which show that our proposed methodology significantly outperforms previous state-of-the-art approaches for gesture generation. Link to code, data and videos: https://github.com/chahuja/aisle
Anthology ID:
2020.findings-emnlp.170
Volume:
Findings of the Association for Computational Linguistics: EMNLP 2020
Month:
November
Year:
2020
Address:
Online
Editors:
Trevor Cohn, Yulan He, Yang Liu
Venue:
Findings
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
1884–1895
Language:
URL:
https://aclanthology.org/2020.findings-emnlp.170
DOI:
10.18653/v1/2020.findings-emnlp.170
Bibkey:
Cite (ACL):
Chaitanya Ahuja, Dong Won Lee, Ryo Ishii, and Louis-Philippe Morency. 2020. No Gestures Left Behind: Learning Relationships between Spoken Language and Freeform Gestures. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1884–1895, Online. Association for Computational Linguistics.
Cite (Informal):
No Gestures Left Behind: Learning Relationships between Spoken Language and Freeform Gestures (Ahuja et al., Findings 2020)
Copy Citation:
PDF:
https://aclanthology.org/2020.findings-emnlp.170.pdf
Optional supplementary material:
 2020.findings-emnlp.170.OptionalSupplementaryMaterial.pdf
Video:
 https://slideslive.com/38940175
Data
PATS