Lucas B. Bulcão Mota

Author directory

2026

Avaliar modelos de linguagem em domínios específicos, como o Direito do Consumidor, é um desafio devido à escassez de recursos padronizados em português brasileiro. Para resolver esse problema, introduzimos o OABench-Consumidor, um benchmark composto exclusivamente por questões de Direito do Consumidor da primeira fase do Exame de Ordem Unificado (OAB), cobrindo as edições de 2010.1 a 2025.2. Este artigo descreve a coleta e a estruturação dos dados. Além disso, a utilidade do benchmark é demonstrada através da avaliação do modelo fundacional Qwen3-8B e de uma versão ajustada. O modelo especializado obteve 64,84% de acurácia (59/91), superando os 51,65% (47/91) do modelo base. O benchmark oferece um recurso direto para avaliar sistemas de Inteligência Artificial nessa subárea do Direito.
The use of textbooks as primary sources of information has increasingly given way to tools based on Large Language Models (LLMs), raising concerns about the reliability of generated answers. This study investigates how different adaptation strategies shape the behavior of small language models in educational Question Answering (QA) tasks in Portuguese. To support this analysis, we built a question-answer dataset derived from an NLP textbook and compared base models and the Retrieval-Augmented Generation (RAG) pipeline with models adapted through supervised fine-tuning and Continued Pretraining. The evaluation relies on questions from the LARI dataset, which has been validated by human specialists, and combines automatic and qualitative assessment procedures. The findings indicate that small models tuned with structured instructional knowledge achieve stronger semantic alignment and produce more pertinent answers in Portuguese educational QA scenarios.