ReSiPC: a Tool for Complex Searches in Parallel Corpora

Antoni Oliver, Bojana Mikelenić


Abstract
In this paper, a tool specifically designed to allow for complex searches in large parallel corpora is presented. The formalism for the queries is very powerful as it uses standard regular expressions that allow for complex queries combining word forms, lemmata and POS-tags. As queries are performed over POS-tags, at least one of the languages in the parallel corpus should be POS-tagged. Searches can be performed in one of the languages or in both languages at the same time. The program is able to POS-tag the corpora using the Freeling analyzer through its Python API. ReSiPC is developed in Python version 3 and it is distributed under a free license (GNU GPL). The tool can be used to provide data for contrastive linguistics research and an example of use in a Spanish-Croatian parallel corpus is presented. ReSiPC is designed for queries in POS-tagged corpora, but it can be easily adapted for querying corpora containing other kinds of information.
Anthology ID:
2020.lrec-1.869
Volume:
Proceedings of the Twelfth Language Resources and Evaluation Conference
Month:
May
Year:
2020
Address:
Marseille, France
Editors:
Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association
Note:
Pages:
7033–7037
Language:
English
URL:
https://aclanthology.org/2020.lrec-1.869
DOI:
Bibkey:
Cite (ACL):
Antoni Oliver and Bojana Mikelenić. 2020. ReSiPC: a Tool for Complex Searches in Parallel Corpora. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 7033–7037, Marseille, France. European Language Resources Association.
Cite (Informal):
ReSiPC: a Tool for Complex Searches in Parallel Corpora (Oliver & Mikelenić, LREC 2020)
Copy Citation:
PDF:
https://aclanthology.org/2020.lrec-1.869.pdf