Training & Quality Assessment of an Optical Character Recognition Model for Northern Haida

Isabell Hubert, Antti Arppe, Jordan Lachler, Eddie A. Santos


Abstract
We are presenting our work on the creation of the first optical character recognition (OCR) model for Northern Haida, also known as Masset or Xaad Kil, a nearly extinct First Nations language spoken in the Haida Gwaii archipelago in British Columbia, Canada. We are addressing the challenges of training an OCR model for a language with an extensive, non-standard Latin character set as follows: (1) We have compared various training approaches and present the results of practical analyses to maximize recognition accuracy and minimize manual labor. An approach using just one or two pages of Source Images directly performed better than the Image Generation approach, and better than models based on three or more pages. Analyses also suggest that a character’s frequency is directly correlated with its recognition accuracy. (2) We present an overview of current OCR accuracy analysis tools available. (3) We have ported the once de-facto standardized OCR accuracy tools to be able to cope with Unicode input. Our work adds to a growing body of research on OCR for particularly challenging character sets, and contributes to creating the largest electronic corpus for this severely endangered language.
Anthology ID:
L16-1514
Volume:
Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)
Month:
May
Year:
2016
Address:
Portorož, Slovenia
Editors:
Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, Helene Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
3227–3234
Language:
URL:
https://aclanthology.org/L16-1514
DOI:
Bibkey:
Cite (ACL):
Isabell Hubert, Antti Arppe, Jordan Lachler, and Eddie A. Santos. 2016. Training & Quality Assessment of an Optical Character Recognition Model for Northern Haida. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), pages 3227–3234, Portorož, Slovenia. European Language Resources Association (ELRA).
Cite (Informal):
Training & Quality Assessment of an Optical Character Recognition Model for Northern Haida (Hubert et al., LREC 2016)
Copy Citation:
PDF:
https://aclanthology.org/L16-1514.pdf