From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model

Marvin Lavechin; Thomas Hueber

doi:10.18653/v1/2025.emnlp-main.1217

From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model

Abstract

Human infants face a formidable challenge in speech acquisition: mapping extremely variable acoustic inputs into appropriate articulatory movements without explicit instruction. We present a computational model that addresses the acoustic-to-articulatory mapping problem through self-supervised learning. Our model comprises a feature extractor that transforms speech into latent representations, an inverse model that maps these representations to articulatory parameters, and a synthesizer that generates speech outputs. Experiments conducted in both single- and multi-speaker settings reveal that intermediate layers of a pre-trained wav2vec 2.0 model provide optimal representations for articulatory learning, significantly outperforming MFCC features. These representations enable our model to learn articulatory trajectories that correlate with human patterns, discriminate between places of articulation, and produce intelligible speech. Critical to successful articulatory learning are representations that balance phonetic discriminability with speaker invariance – precisely the characteristics of self-supervised representation learning models. Our findings provide computational evidence consistent with developmental theories proposing that perceptual learning of phonetic categories guides articulatory development, offering insights into how infants might acquire speech production capabilities despite the complex mapping problem they face.

Anthology ID:: 2025.emnlp-main.1217
Volume:: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 23852–23863
Language:
URL:: https://aclanthology.org/2025.emnlp-main.1217/
DOI:: 10.18653/v1/2025.emnlp-main.1217
Bibkey:
Cite (ACL):: Marvin Lavechin and Thomas Hueber. 2025. From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 23852–23863, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model (Lavechin & Hueber, EMNLP 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.emnlp-main.1217.pdf
Checklist:: 2025.emnlp-main.1217.checklist.pdf

PDF Cite Search Checklist Fix data