A Benchmark Dataset and Comparative Evaluation of Phonemized and Romanized Urdu for Text-to-Speech

M Kaab Bin Shahid, Muhammed Izharuddin


Abstract
Text-to-Speech (TTS) system for the Urdu language presents significant challenges, primarily due to the scarcity of high-quality datasets and an insufficient focus on modeling pronunciation. Urdu is spoken by 250 million people worldwide, but its research on computational linguistics remains underrepresented. In this paper, we introduce URDUTTS, a comprehensive and publicly available Urdu TTS dataset containing 89 hours of studio-quality speech, with accompanying transcriptions in three formats: Urdu Script, Phonemized Script, and Romanized Script. The dataset includes both mono-speaker and multi-speaker configurations. As Urdu relies heavily on phonetic features, accurate pronunciation is highly essential for the language. Therefore, we benchmark our dataset using VITS and GlowTTS models to compare the widely used Romanized script format with the Phonemized representation. To make the evaluation highly comprehensive, we combined both objective and subjective evaluation strategies. For objective evaluation, Mel-Cepstral Distortion (MCD with Plain, Dynamic Time-Warping, and Slope-Limitation variants), Signal-to-Noise Ratio (SNR), Word Error Rate (WER), and Character Error Rate (CER) were taken. Subjective evaluation was governed by Mean Opinion Score (MOS) ratings from 40 native speakers. Results show that using VITS and GlowTTS with Phonemized transcriptions performs significantly better than Romanized ones, with an improvement of 9.6% and 26.5% in MOS. The data and code are available at github.com/KAABSHAHID/URDUTTS.
Anthology ID:
2026.lrec-1.859
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
10982–10993
Language:
External URL:
https://lrec.elra.info/lrec2026-main-859
DOI:
10.63317/2avnr98mgbre
Bibkey:
Cite (ACL):
M Kaab Bin Shahid and Muhammed Izharuddin. 2026. A Benchmark Dataset and Comparative Evaluation of Phonemized and Romanized Urdu for Text-to-Speech. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 10982–10993, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
A Benchmark Dataset and Comparative Evaluation of Phonemized and Romanized Urdu for Text-to-Speech (Shahid & Izharuddin, LREC 2026)
Copy Citation: