Miguel Carpi
Author directoryAlso published as: Miguel M. Carpi
2026
Exploring Hybrid Pre-training for Automatic Essay Scoring
Miguel M. Carpi | Igor C. Silveira | Denis D. Mauá | Marcelo Finger
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Miguel M. Carpi | Igor C. Silveira | Denis D. Mauá | Marcelo Finger
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Alternatives to Large Language Models have been proposed to develop smaller models that require substantially less training data. In this paper, we propose a monolingual (Portuguese) Hybrid Transformer model trained with only 260M words, whose size is comparable to that of small Encoder-based models. After pre-training, we fine-tune our model and compare it against 11 existing models on the AES-ENEM dataset — an Automatic Essay Scoring benchmark in which models are required to evaluate five distinct textual dimensions. Our experiments demonstrate that our model is always competitive with the (bigger) best available model, despite its smaller scale.
2024
Analysing and Validating Language Complexity Metrics Across South American Indigenous Languages
Felipe Serras | Miguel Carpi | Matheus Branco | Marcelo Finger
Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics
Felipe Serras | Miguel Carpi | Matheus Branco | Marcelo Finger
Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics
Language complexity is an emerging concept critical for NLP and for quantitative and cognitive approaches to linguistics. In this work, we evaluate the behavior of a set of compression-based language complexity metrics when applied to a large set of native South American languages. Our goal is to validate the desirable properties of such metrics against a more diverse set of languages, guaranteeing the universality of the techniques developed on the basis of this type of theoretical artifact. Our analysis confirmed with statistical confidence most propositions about the metrics studied, affirming their robustness, despite showing less stability than when the same metrics were applied to Indo-European languages. We also observed that the trade-off between morphological and syntactic complexities is strongly related to language phylogeny.