Exploring Hybrid Pre-training for Automatic Essay Scoring

Miguel M. Carpi, Igor C. Silveira, Denis D. Mauá, Marcelo Finger


Abstract
Alternatives to Large Language Models have been proposed to develop smaller models that require substantially less training data. In this paper, we propose a monolingual (Portuguese) Hybrid Transformer model trained with only 260M words, whose size is comparable to that of small Encoder-based models. After pre-training, we fine-tune our model and compare it against 11 existing models on the AES-ENEM dataset — an Automatic Essay Scoring benchmark in which models are required to evaluate five distinct textual dimensions. Our experiments demonstrate that our model is always competitive with the (bigger) best available model, despite its smaller scale.
Anthology ID:
2026.stil-1.6
Volume:
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Month:
October
Year:
2026
Address:
Cuiabá, Mato Grosso, Brazil
Editors:
Bryan Khelven da Silva Barbosa, Aline Paes, Ariani Di Felippo
Venue:
STIL
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
62–77
Language:
URL:
https://aclanthology.org/2026.stil-1.6/
DOI:
10.5753/stil.2026.26589
Bibkey:
Cite (ACL):
Miguel M. Carpi, Igor C. Silveira, Denis D. Mauá, and Marcelo Finger. 2026. Exploring Hybrid Pre-training for Automatic Essay Scoring. In Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology, pages 62–77, Cuiabá, Mato Grosso, Brazil. Association for Computational Linguistics.
Cite (Informal):
Exploring Hybrid Pre-training for Automatic Essay Scoring (Carpi et al., STIL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.stil-1.6.pdf