TantaArabNLP at KSAA-2026 Task 2: Adapting CATT-Whisper for Arabic Speech Dictation with Automatic Diacritization

Nada Adel Esmaeil, Reda M. Elbasiony, Mohamed T. Faheem


Abstract
We present our submission to the KSAA-2026 Shared Task (Subtask 2): Automatic Diacritization of Speech Dictation. Building upon the CATT-Whisper multimodal architecture, which fuses representations from a pre-trained CATT text encoder and the Whisper speech encoder, we fine-tune the model end-to-end on the official shared task training data. To further enhance performance on speech-dictated Arabic text, we apply careful post-processing to the model outputs. Our best submission achieves a Diacritic Error Rate (DER) of 7.04, a Word Error Rate (WER) of 24.39, and a Sentence Error Rate (SER) of 71.65 on the hidden test set, securing 2nd place in the competition. These results demonstrate the effectiveness of adapting a strong multimodal baseline to the speech-aware diacritization setting and highlight the value of task-specific fine-tuning and output refinement for bridging the gap between spoken transcripts and fully diacritized Arabic text.
Anthology ID:
2026.osact-1.30
Volume:
The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Hend Al-Khalifa, Mo El-Haj, Saad Ezzini
Venues:
OSACT | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
229–233
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-osact-30
DOI:
10.63317/46cm97fcekow
Bibkey:
Cite (ACL):
Nada Adel Esmaeil, Reda M. Elbasiony, and Mohamed T. Faheem. 2026. TantaArabNLP at KSAA-2026 Task 2: Adapting CATT-Whisper for Arabic Speech Dictation with Automatic Diacritization. In The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks, pages 229–233, Palma, Mallorca (Spain). Association for Computational Linguistics.
Cite (Informal):
TantaArabNLP at KSAA-2026 Task 2: Adapting CATT-Whisper for Arabic Speech Dictation with Automatic Diacritization (Esmaeil et al., OSACT 2026)
Copy Citation: