Combining Multiple Corpora for Readability Assessment for People with Cognitive Disabilities

Victoria Yaneva; Constantin Orasan; Richard Evans; Omid Rohanian

doi:10.18653/v1/W17-5013

Combining Multiple Corpora for Readability Assessment for People with Cognitive Disabilities

Victoria Yaneva, Constantin Orăsan, Richard Evans, Omid Rohanian

Abstract

Given the lack of large user-evaluated corpora in disability-related NLP research (e.g. text simplification or readability assessment for people with cognitive disabilities), the question of choosing suitable training data for NLP models is not straightforward. The use of large generic corpora may be problematic because such data may not reflect the needs of the target population. The use of the available user-evaluated corpora may be problematic because these datasets are not large enough to be used as training data. In this paper we explore a third approach, in which a large generic corpus is combined with a smaller population-specific corpus to train a classifier which is evaluated using two sets of unseen user-evaluated data. One of these sets, the ASD Comprehension corpus, is developed for the purposes of this study and made freely available. We explore the effects of the size and type of the training data used on the performance of the classifiers, and the effects of the type of the unseen test datasets on the classification performance.

Anthology ID:: W17-5013
Volume:: Proceedings of the 12th Workshop on Innovative Use of NLP for Building Educational Applications
Month:: September
Year:: 2017
Address:: Copenhagen, Denmark
Editors:: Joel Tetreault, Jill Burstein, Claudia Leacock, Helen Yannakoudakis
Venue:: BEA
SIG:: SIGEDU
Publisher:: Association for Computational Linguistics
Note:
Pages:: 121–132
Language:
URL:: https://aclanthology.org/W17-5013/
DOI:: 10.18653/v1/W17-5013
Bibkey:
Cite (ACL):: Victoria Yaneva, Constantin Orăsan, Richard Evans, and Omid Rohanian. 2017. Combining Multiple Corpora for Readability Assessment for People with Cognitive Disabilities. In Proceedings of the 12th Workshop on Innovative Use of NLP for Building Educational Applications, pages 121–132, Copenhagen, Denmark. Association for Computational Linguistics.
Cite (Informal):: Combining Multiple Corpora for Readability Assessment for People with Cognitive Disabilities (Yaneva et al., BEA 2017)
Copy Citation:
PDF:: https://aclanthology.org/W17-5013.pdf

PDF Cite Search Fix data