Textbook-Enriched Training for Language Models: Boosting Answer Quality in Specialized Contexts

Lucas B. Bulcão Mota, Larrissa Dantas, Daniela Barreiro Claro, Aline Paes, Claudia Freitas, Marlo Souza, Helena Caseli, Livy Real


Abstract
The use of textbooks as primary sources of information has increasingly given way to tools based on Large Language Models (LLMs), raising concerns about the reliability of generated answers. This study investigates how different adaptation strategies shape the behavior of small language models in educational Question Answering (QA) tasks in Portuguese. To support this analysis, we built a question-answer dataset derived from an NLP textbook and compared base models and the Retrieval-Augmented Generation (RAG) pipeline with models adapted through supervised fine-tuning and Continued Pretraining. The evaluation relies on questions from the LARI dataset, which has been validated by human specialists, and combines automatic and qualitative assessment procedures. The findings indicate that small models tuned with structured instructional knowledge achieve stronger semantic alignment and produce more pertinent answers in Portuguese educational QA scenarios.
Anthology ID:
2026.stil-1.23
Volume:
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Month:
October
Year:
2026
Address:
Cuiabá, Mato Grosso, Brazil
Editors:
Bryan Khelven da Silva Barbosa, Aline Paes, Ariani Di Felippo
Venue:
STIL
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
270–284
Language:
URL:
https://aclanthology.org/2026.stil-1.23/
DOI:
10.5753/stil.2026.26566
Bibkey:
Cite (ACL):
Lucas B. Bulcão Mota, Larrissa Dantas, Daniela Barreiro Claro, Aline Paes, Claudia Freitas, Marlo Souza, Helena Caseli, and Livy Real. 2026. Textbook-Enriched Training for Language Models: Boosting Answer Quality in Specialized Contexts. In Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology, pages 270–284, Cuiabá, Mato Grosso, Brazil. Association for Computational Linguistics.
Cite (Informal):
Textbook-Enriched Training for Language Models: Boosting Answer Quality in Specialized Contexts (Mota et al., STIL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.stil-1.23.pdf