Selection of Japanese-English Equivalents by Integrating High-quality Corpora and Huge Amounts of Web Data

Qing Ma; Koichi Nakao; Masaki Murata; Hitoshi Isahara

Selection of Japanese-English Equivalents by Integrating High-quality Corpora and Huge Amounts of Web Data

Qing Ma, Koichi Nakao, Masaki Murata, Hitoshi Isahara

Abstract

As a first step to developing systems that enable non-native speakers to output near-perfect English sentences for given mixed English-Japanese sentences, we propose new approaches for selecting English equivalents by using the number of hits for various contexts in large English corpora. As the large English corpora, we not only used the huge amounts of Web data but also the manually compiled large, high-quality English corpora. Using high-quality corpora enables us to accurately select equivalents, and using huge amounts of Web data enables us to resolve the problem of the shortage of hits that normally occurs when using only high-quality corpora. The types and lengths of contexts used to select equivalents are variable and optimally determined according to the number of hits in the corpora, so that performance can be further refined. Computer experiments showed that the precision of our methods was much higher than that of the existing methods for equivalent selection.

Anthology ID:: L08-1570
Volume:: Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)
Month:: May
Year:: 2008
Address:: Marrakech, Morocco
Editors:: Nicoletta Calzolari, Khalid Choukri, Bente Maegaard, Joseph Mariani, Jan Odijk, Stelios Piperidis, Daniel Tapias
Venue:: LREC
SIG:
Publisher:: European Language Resources Association (ELRA)
Note:
Pages:
Language:
URL:: http://www.lrec-conf.org/proceedings/lrec2008/pdf/107_paper.pdf
DOI:
Bibkey:
Cite (ACL):: Qing Ma, Koichi Nakao, Masaki Murata, and Hitoshi Isahara. 2008. Selection of Japanese-English Equivalents by Integrating High-quality Corpora and Huge Amounts of Web Data. In Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08), Marrakech, Morocco. European Language Resources Association (ELRA).
Cite (Informal):: Selection of Japanese-English Equivalents by Integrating High-quality Corpora and Huge Amounts of Web Data (Ma et al., LREC 2008)
Copy Citation:
PDF:: http://www.lrec-conf.org/proceedings/lrec2008/pdf/107_paper.pdf

PDF Cite Search Fix data