Robust clause boundary identification for corpus annotation

Heiki-Jaan Kaalep, Kadri Muischnek


Abstract
The paper describes a rule-based system for tagging clause boundaries, implemented for annotating the Estonian Reference Corpus of the University of Tartu, a collection of written texts containing ca 245 million running words and available for querying via Keeleveeb language portal. The system needs information about parts of speech and grammatical categories coded in the word-forms, i.e. it takes morphologically annotated text as input, but requires no information about the syntactic structure of the sentence. Among the strong points of our system we should mention identifying parenthesis and embedded clauses, i.e. clauses that are inserted into another clause dividing it into two separate parts in the linear text, for example a relative clause following its head noun. That enables a corpus query system to unite the otherwise divided clause, a feature that usually presupposes full parsing. The overall precision of the system is 95% and the recall is 96%. If “ordinary” clause boundary detection and parenthesis and embedded clause boundary detection are evaluated separately, then one can say that detecting an “ordinary” clause boundary (recall 98%, precision 96%) is an easier task than detecting an embedded clause (recall 79%, precision 100%).
Anthology ID:
L12-1083
Volume:
Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
Month:
May
Year:
2012
Address:
Istanbul, Turkey
Editors:
Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
1632–1636
Language:
URL:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/229_Paper.pdf
DOI:
Bibkey:
Cite (ACL):
Heiki-Jaan Kaalep and Kadri Muischnek. 2012. Robust clause boundary identification for corpus annotation. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12), pages 1632–1636, Istanbul, Turkey. European Language Resources Association (ELRA).
Cite (Informal):
Robust clause boundary identification for corpus annotation (Kaalep & Muischnek, LREC 2012)
Copy Citation:
PDF:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/229_Paper.pdf