Claudia Freitas
Author directoryOther people with similar names: Cláudia Freitas
2026
Textbook-Enriched Training for Language Models: Boosting Answer Quality in Specialized Contexts
Lucas B. Bulcão Mota | Larrissa Dantas | Daniela Barreiro Claro | Aline Paes | Claudia Freitas | Marlo Souza | Helena Caseli | Livy Real
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Lucas B. Bulcão Mota | Larrissa Dantas | Daniela Barreiro Claro | Aline Paes | Claudia Freitas | Marlo Souza | Helena Caseli | Livy Real
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
The use of textbooks as primary sources of information has increasingly given way to tools based on Large Language Models (LLMs), raising concerns about the reliability of generated answers. This study investigates how different adaptation strategies shape the behavior of small language models in educational Question Answering (QA) tasks in Portuguese. To support this analysis, we built a question-answer dataset derived from an NLP textbook and compared base models and the Retrieval-Augmented Generation (RAG) pipeline with models adapted through supervised fine-tuning and Continued Pretraining. The evaluation relies on questions from the LARI dataset, which has been validated by human specialists, and combines automatic and qualitative assessment procedures. The findings indicate that small models tuned with structured instructional knowledge achieve stronger semantic alignment and produce more pertinent answers in Portuguese educational QA scenarios.