Towards an Automatic Classification of Illustrative Examples in a Large Japanese-French Dictionary Obtained by OCR

Christian Boitet; Mathieu Mangeot; Mutsuko Tomokiyo

Towards an Automatic Classification of Illustrative Examples in a Large Japanese-French Dictionary Obtained by OCR

Christian Boitet, Mathieu Mangeot, Mutsuko Tomokiyo

Abstract

We work on improving the Cesselin, a large and open source Japanese-French bilingual dictionary digitalized by OCR, available on the web, and contributively improvable online. Labelling its examples (about 226000) would significantly enhance their usefulness for language learners. Examples are proverbs, idiomatic constructions, normal usage examples, and, for nouns, phrases containing a quantifier. Proverbs are easy to spot, but not examples of other types. To find a method for automatically or at least semi-automatically annotating them, we have studied many entries, and hypothesized that the degree of lexical similarity between results of MT into a third language might give good cues. To confirm that hypothesis, we sampled 500 examples and used Google Translate to translate into English their Japanese expressions and their French translations. The hypothesis holds well, in particular for distinguishing examples of normal usage from idiomatic examples. Finally, we propose a detailed annotation procedure and discuss its future automatization.

Anthology ID:: W18-3815
Volume:: Proceedings of the First Workshop on Linguistic Resources for Natural Language Processing
Month:: August
Year:: 2018
Address:: Santa Fe, New Mexico, USA
Editors:: Peter Machonis, Anabela Barreiro, Kristina Kocijan, Max Silberztein
Venue:: LR4NLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 112–121
Language:
URL:: https://aclanthology.org/W18-3815/
DOI:
Bibkey:
Cite (ACL):: Christian Boitet, Mathieu Mangeot, and Mutsuko Tomokiyo. 2018. Towards an Automatic Classification of Illustrative Examples in a Large Japanese-French Dictionary Obtained by OCR. In Proceedings of the First Workshop on Linguistic Resources for Natural Language Processing, pages 112–121, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
Cite (Informal):: Towards an Automatic Classification of Illustrative Examples in a Large Japanese-French Dictionary Obtained by OCR (Boitet et al., LR4NLP 2018)
Copy Citation:
PDF:: https://aclanthology.org/W18-3815.pdf

PDF Cite Search Fix data