Felipe Serras
Author directoryAlso published as: Felipe Ribas Serras
2026
Compression-Based Linguistic Complexity Metrics in Automatic Essay Scoring
Felipe Ribas Serras | Igor Cataneo Silveira | Denis Deratani Mauá | Marcelo Finger
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Felipe Ribas Serras | Igor Cataneo Silveira | Denis Deratani Mauá | Marcelo Finger
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Compression-based linguistic complexity metrics enable cross-linguistic comparison without prior annotation. Their sensitivity to variation across languages and Portuguese registers highlights their applicability in NLP tasks. This study investigates their use as readability proxies and complementary features in Automatic Essay Scoring. We analyze how these metrics capture variation in essay quality across traits, genres, and educational levels in Brazilian Portuguese. In addition, we evaluate their sensitivity to differences between humanand AI-generated essays. Our results suggest that complexity metrics are effective (i) in differentiating educational levels, (ii) in detecting whether they were written by humans and (iii) as predictors of essay quality.
Compression-based Language Complexity under Register Variation in Portuguese
Felipe Ribas Serras | Marcelo Finger
Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1
Felipe Ribas Serras | Marcelo Finger
Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1
Compression-based language complexity metrics show promise as holistic parameters for measuring linguistic complexity across intra- and cross-linguistic scenarios. Yet, their sensitivity to specific forms of linguistic variation requires further experimental validation. We examine the sensitivity of this metric family to register variation in Portuguese, a phenomenon already established for English. We refine the validation process found in previous literature by introducing a more granular statistical analysis to evaluate both the individual and joint sensitivity of these metrics to register variation at the sentence level. Our results confirm they are highly sensitive to functional variation in Portuguese, exhibiting the same structural morphosyntactic trade-off consistent with that observed in English and in cross-linguistic studies.
2024
Exploring Computational Discernibility of Discourse Domains in Brazilian Portuguese within the Carolina Corpus
Felipe Ribas Serras | Mariana Sturzeneker | Miguel de Mello Carpi | Mayara Feliciano Palma | Maria Clara Ramos Morales Crespo | Aline Silva Costa | Vanessa Martins Do Monte | Cristiane Namiuti | Maria Clara Paixão de Souza | Marcelo Finger
Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1
Felipe Ribas Serras | Mariana Sturzeneker | Miguel de Mello Carpi | Mayara Feliciano Palma | Maria Clara Ramos Morales Crespo | Aline Silva Costa | Vanessa Martins Do Monte | Cristiane Namiuti | Maria Clara Paixão de Souza | Marcelo Finger
Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1
Analysing and Validating Language Complexity Metrics Across South American Indigenous Languages
Felipe Serras | Miguel Carpi | Matheus Branco | Marcelo Finger
Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics
Felipe Serras | Miguel Carpi | Matheus Branco | Marcelo Finger
Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics
Language complexity is an emerging concept critical for NLP and for quantitative and cognitive approaches to linguistics. In this work, we evaluate the behavior of a set of compression-based language complexity metrics when applied to a large set of native South American languages. Our goal is to validate the desirable properties of such metrics against a more diverse set of languages, guaranteeing the universality of the techniques developed on the basis of this type of theoretical artifact. Our analysis confirmed with statistical confidence most propositions about the metrics studied, affirming their robustness, despite showing less stability than when the same metrics were applied to Indo-European languages. We also observed that the trade-off between morphological and syntactic complexities is strongly related to language phylogeny.