From Press to Pixels: Evolving Urdu Text Recognition

Samee Arif, Sualeha Farid


Abstract
This paper presents a comparative analysis of Large Language Models (LLMs) and traditional Optical Character Recognition (OCR) systems on Urdu newspapers, addressing challenges posed by complex multi-column layouts, low-resolution scans, and the stylistic variability of the Nastaliq script. To handle these challenges, we fine-tune YOLOv11x models for article- and column-level text block extraction and train a SwinIR-based super-resolution module that enhances image quality for downstream text recognition, improving accuracy by an average of 50%. We further introduce the Urdu Newspaper Benchmark (UNB), a manually annotated dataset for Urdu OCR comprising 829 paragraph images with a total of 9,982 sentences. Using UNB and the OpenITI corpus, we conduct a systematic comparison between traditional CNN+RNN-based OCR systems and modern LLMs, presenting detailed insertion, deletion, and substitution error analyses alongside character-level confusion patterns. We find that Gemini-2.5-Pro achieves the best performance on UNB (WER 0.133), while fine-tuning GPT-4o on just 500 in-domain samples yields a 6.13% absolute WER improvement, demonstrating the adaptability of LLMs to low-resource, morphologically complex scripts like Urdu. The UNB dataset and fine-tuned models are publicly available at https://github.com/paper-seven/UrduOCR.
Anthology ID:
2026.lrec-1.235
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
3013–3021
Language:
External URL:
https://lrec.elra.info/lrec2026-main-235
DOI:
10.63317/4drrpn75kzpm
Bibkey:
Cite (ACL):
Samee Arif and Sualeha Farid. 2026. From Press to Pixels: Evolving Urdu Text Recognition. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 3013–3021, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
From Press to Pixels: Evolving Urdu Text Recognition (Arif & Farid, LREC 2026)
Copy Citation: