Bootstrapping the lexicon building process for machine translation between ‘new’ languages

Ruvan Weerasinghe

Bootstrapping the lexicon building process for machine translation between ‘new’ languages

Abstract

The cumulative effort over the past few decades that have gone into developing linguistic resources for tasks ranging from machine readable dictionaries to translation systems is enormous. Such effort is prohibitively expensive for languages outside the (largely) European family. The possibility of building such resources automatically by accessing electronic corpora of such languages are therefore of great interest to those involved in studying these ‘new’ - ‘lesser known’ languages. The main stumbling block to applying these data driven techniques directly is that most of them require large corpora rarely available for such ‘new’ languages. This paper describes an attempt at setting up a bootstrapping agenda to exploit the scarce corpus resources that may be available at the outset to a researcher concerned with such languages. In particular it reports on results of an experiment to use state-of-the-art data-driven techniques for building linguistic resources for Sinhala - a non-European language with virtually no electronic resources.

Anthology ID:: 2002.amta-papers.18
Volume:: Proceedings of the 5th Conference of the Association for Machine Translation in the Americas: Technical Papers
Month:: October 8-12
Year:: 2002
Address:: Tiburon, USA
Editor:: Stephen D. Richardson
Venue:: AMTA
SIG:
Publisher:: Springer
Note:
Pages:: 177–186
Language:
URL:: https://link.springer.com/chapter/10.1007/3-540-45820-4_18
DOI:
Bibkey:
Cite (ACL):: Ruvan Weerasinghe. 2002. Bootstrapping the lexicon building process for machine translation between ‘new’ languages. In Proceedings of the 5th Conference of the Association for Machine Translation in the Americas: Technical Papers, pages 177–186, Tiburon, USA. Springer.
Cite (Informal):: Bootstrapping the lexicon building process for machine translation between ‘new’ languages (Weerasinghe, AMTA 2002)
Copy Citation:
PDF:: https://link.springer.com/chapter/10.1007/3-540-45820-4_18

PDF Cite Search Fix data