G&P2P: A Multi-Source Approach to Grapheme-to-Phoneme Conversion

Chun-Yi Jerry Peng


Abstract
Grapheme-to-phoneme (G2P) conversion plays a central role in speech technologies. This paper introduces G&P2P, a multi-source framework that integrates multiple pronunciation dictionaries to enhance G2P modeling. We evaluate both expert-curated and crowd-sourced resources using attentive LSTM, pointer-generator LSTM, and transformer architectures. Results indicate that combining high-quality expert dictionaries yields substantial improvements, achieving an 11.26-point absolute (22% relative) reduction in word error rate. In contrast, incorporating noisy crowd-sourced resources may degrade performance. Statistical analyses further suggest that dataset quality exerts a greater influence on outcomes than the choice of fusion strategy, offering practical guidance for the design of multi-source G2P systems.
Anthology ID:
2026.cawl-1.9
Volume:
Proceedings of the Third Workshop on Computation and Written Language (CAWL 2026) @ LREC 2026
Month:
June
Year:
2026
Address:
Palma de Mallorca, Spain
Editor:
Kyle Gorman
Venues:
CAWL | WS
SIG:
SIGWrit
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
89–94
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-cawl-09
DOI:
10.63317/4f2x3fda6jj7
Bibkey:
Cite (ACL):
Chun-Yi Jerry Peng. 2026. G&P2P: A Multi-Source Approach to Grapheme-to-Phoneme Conversion. In Proceedings of the Third Workshop on Computation and Written Language (CAWL 2026) @ LREC 2026, pages 89–94, Palma de Mallorca, Spain. ELRA Language Resources Association (ELRA).
Cite (Informal):
G&P2P: A Multi-Source Approach to Grapheme-to-Phoneme Conversion (Peng, CAWL 2026)
Copy Citation: