Hypergraph Modelization of a Syntactically Annotated English Wikipedia Dump

Edmundo Pavel Soriano Morales, Julien Ah-Pine, Sabine Loudcher


Abstract
Wikipedia, the well known internet encyclopedia, is nowadays a widely used source of information. To leverage its rich information, already parsed versions of Wikipedia have been proposed. We present an annotated dump of the English Wikipedia. This dump draws upon previously released Wikipedia parsed dumps. Still, we head in a different direction. In this parse we focus more into the syntactical characteristics of words: aside from the classical Part-of-Speech (PoS) tags and dependency parsing relations, we provide the full constituent parse branch for each word in a succinct way. Additionally, we propose a hypergraph network representation of the extracted linguistic information. The proposed modelization aims to take advantage of the information stocked within our parsed Wikipedia dump. We hope that by releasing these resources, researchers from the concerned communities will have a ready-to-experiment Wikipedia corpus to compare and distribute their work. We render public our parsed Wikipedia dump as well as the tool (and its source code) used to perform the parse. The hypergraph network and its related metadata is also distributed.
Anthology ID:
L16-1372
Volume:
Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)
Month:
May
Year:
2016
Address:
Portorož, Slovenia
Editors:
Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, Helene Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
2348–2353
Language:
URL:
https://aclanthology.org/L16-1372
DOI:
Bibkey:
Cite (ACL):
Edmundo Pavel Soriano Morales, Julien Ah-Pine, and Sabine Loudcher. 2016. Hypergraph Modelization of a Syntactically Annotated English Wikipedia Dump. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), pages 2348–2353, Portorož, Slovenia. European Language Resources Association (ELRA).
Cite (Informal):
Hypergraph Modelization of a Syntactically Annotated English Wikipedia Dump (Morales et al., LREC 2016)
Copy Citation:
PDF:
https://aclanthology.org/L16-1372.pdf