Oblevit at AR-MS NAKBA NLP 2026 Subtask 2: Hybrid CNNBiLSTMCTC Framework with Linguistic Refinement for Arabic Handwritten Manuscript Recognition

Reem Juhaysh, Abuelgasim Sami Abusonoun, Sara Ayad


Abstract
Arabic handwritten manuscript recognition is challenging due to the cursive nature of the script, dot ambiguity, and document degradation. In this work, we propose an end-to-end OCR system based on a CNN–BiLSTM–CTC architecture. The model extracts visual features, captures sequential dependencies, and performs alignment-free training. Arabic-specific decoding and post-processing techniques are applied to reduce character and spacing errors. Experimental results show competitive performance in recognizing complex handwritten Arabic text.
Anthology ID:
2026.nakbanlp-1.35
Volume:
Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Mustafa Jarrar, Mo El-Haj, Amal Haddad, Serin Atiani, Shadi Abudalfa, Terry Regier, Paul Rayson, Khalil Sima’an, Camille Mansour
Venues:
NakbaNLP | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
234–238
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-nakbanlp-35
DOI:
10.63317/5nwur655ha5k
Bibkey:
Cite (ACL):
Reem Juhaysh, Abuelgasim Sami Abusonoun, and Sara Ayad. 2026. Oblevit at AR-MS NAKBA NLP 2026 Subtask 2: Hybrid CNN–BiLSTM–CTC Framework with Linguistic Refinement for Arabic Handwritten Manuscript Recognition. In Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026, pages 234–238, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
Oblevit at AR-MS NAKBA NLP 2026 Subtask 2: Hybrid CNN–BiLSTM–CTC Framework with Linguistic Refinement for Arabic Handwritten Manuscript Recognition (Juhaysh et al., NakbaNLP 2026)
Copy Citation: