Extracting Interlinear Glossed Text from LaTeX Documents

Mathias Schenner, Sebastian Nordhoff


Abstract
We present texigt, a command-line tool for the extraction of structured linguistic data from LaTeX source documents, and a language resource that has been generated using this tool: a corpus of interlinear glossed text (IGT) extracted from open access books published by Language Science Press. Extracted examples are represented in a simple XML format that is easy to process and can be used to validate certain aspects of interlinear glossed text. The main challenge involved is the parsing of TeX and LaTeX documents. We review why this task is impossible in general and how the texhs Haskell library uses a layered architecture and selective early evaluation (expansion) during lexing and parsing in order to provide access to structured representations of LaTeX documents at several levels. In particular, its parsing modules generate an abstract syntax tree for LaTeX documents after expansion of all user-defined macros and lexer-level commands that serves as an ideal interface for the extraction of interlinear glossed text by texigt. This architecture can easily be adapted to extract other types of linguistic data structures from LaTeX source documents.
Anthology ID:
L16-1638
Volume:
Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)
Month:
May
Year:
2016
Address:
Portorož, Slovenia
Editors:
Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, Helene Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
4044–4048
Language:
URL:
https://aclanthology.org/L16-1638
DOI:
Bibkey:
Cite (ACL):
Mathias Schenner and Sebastian Nordhoff. 2016. Extracting Interlinear Glossed Text from LaTeX Documents. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), pages 4044–4048, Portorož, Slovenia. European Language Resources Association (ELRA).
Cite (Informal):
Extracting Interlinear Glossed Text from LaTeX Documents (Schenner & Nordhoff, LREC 2016)
Copy Citation:
PDF:
https://aclanthology.org/L16-1638.pdf