Hierarchical Representation Alignment Learning of Diffusion Transformers for Neural Audio Codec

Sang-Hoon Lee, Ha-Yeong Choi


Abstract
Despite recent progress in diffusion and conditional flow matching (CFM) models for low-resolution domains such as latent representations, their application to high-resolution data like raw waveform signals remains underexplored. Generative adversarial networks (GANs) have been the dominant approach in neural vocoder and neural audio codecs for realistic waveform generation. However, under low-bitrate conditions, these models suffer from degraded performance due to information loss caused by heavy compression and quantization, often resulting in mispronunciations. To address the aforementioned problem, we first leverage CFM to iteratively generate raw waveform in an extremely low-bitrate scenario. We then introduce hierarchical representation alignment learning (REPA-H) to enable efficient and robust CFM training. Furthermore, we propose dense vector quantization (DVQ), a novel factorized quantization method using a single quantizer. Our model, FlowTokenizer, outperforms state-of-the-art neural audio codecs in audio quality and semantic intelligibility under low-bitrate conditions, using only 25 tokens per second for 24 kHz waveform generation.
Anthology ID:
2026.findings-acl.1622
Volume:
Findings of the Association for Computational Linguistics: ACL 2026
Month:
July
Year:
2026
Address:
San Diego, California, United States
Editors:
Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:
Findings
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
32410–32426
Language:
URL:
https://aclanthology.org/2026.findings-acl.1622/
DOI:
10.18653/v1/2026.findings-acl.1622
Bibkey:
Cite (ACL):
Sang-Hoon Lee and Ha-Yeong Choi. 2026. Hierarchical Representation Alignment Learning of Diffusion Transformers for Neural Audio Codec. In Findings of the Association for Computational Linguistics: ACL 2026, pages 32410–32426, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):
Hierarchical Representation Alignment Learning of Diffusion Transformers for Neural Audio Codec (Lee & Choi, Findings 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.findings-acl.1622.pdf
Checklist:
 2026.findings-acl.1622.checklist.pdf