Extended Context at the Introduction of Complex Vocabulary in Abridged Literary Texts

Iglika Nikolova-Stoupak, Eva Schaeffer-Lacroix, Gaël Lejeune


Abstract
Psycholinguistics speaks of a fine-tuning process used by parents as they address children, in which complex vocabulary is introduced with additional context (Leung et al., 2021). This somewhat counterintuitive lengthening of text in order to aid one’s interlocutor in the process of language acquisition also comes in accord with Harris (1988)’s notion that for every complex sentence, there is an equivalent longer (non-contracted) yet simpler one that contains the same amount of information. Within the proposed work, a corpus of eight renowned literary works (e.g. Alice’s Adventures in Wonderland, The Adventures of Tom Sawyer, Les Misérables) in four distinct languages (English, French, Russian and Spanish) is gathered: both the original (or translated) versions and up to four abridged versions for various audiences (e.g. children of a defined age or foreign language learners of a defined level) are present. The contexts of the first appearance of complex words (as determined based on word frequency) in pairs of original and abridged works are compared, and the cases in which the abridged texts offer longer context are investigated. The discovered transformations are consequently classified into three separate categories: addition of vocabulary items from the same lexical field as the complex word, simplification of grammar and insertion of a definition. Context extensions are then statistically analysed as associated with different languages and reader audiences.
Anthology ID:
2024.clib-1.17
Volume:
Proceedings of the Sixth International Conference on Computational Linguistics in Bulgaria (CLIB 2024)
Month:
September
Year:
2024
Address:
Sofia, Bulgaria
Venue:
CLIB
SIG:
Publisher:
Department of Computational Linguistics, Institute for Bulgarian Language, Bulgarian Academy of Sciences
Note:
Pages:
166–177
Language:
URL:
https://aclanthology.org/2024.clib-1.17
DOI:
Bibkey:
Cite (ACL):
Iglika Nikolova-Stoupak, Eva Schaeffer-Lacroix, and Gaël Lejeune. 2024. Extended Context at the Introduction of Complex Vocabulary in Abridged Literary Texts. In Proceedings of the Sixth International Conference on Computational Linguistics in Bulgaria (CLIB 2024), pages 166–177, Sofia, Bulgaria. Department of Computational Linguistics, Institute for Bulgarian Language, Bulgarian Academy of Sciences.
Cite (Informal):
Extended Context at the Introduction of Complex Vocabulary in Abridged Literary Texts (Nikolova-Stoupak et al., CLIB 2024)
Copy Citation:
PDF:
https://aclanthology.org/2024.clib-1.17.pdf