Larrissa Dantas

Author directory

2026

A simplificação de sentenças frequentemente altera a estrutura textual enquanto preserva o significado, tornando a fidelidade semântica difícil de avaliar. Neste trabalho, investigamos se representações baseadas em grafos derivadas de Extração Aberta de Informação (Open IE) podem capturar a preservação semântica em diferentes níveis de simplificação. Utilizando o corpus PorSimplesSent, extraímos triplas relacionais de sentenças originais e simplificadas para construir grafos semânticos comparáveis. Os resultados fornecem indícios exploratórios de que a análise baseada em grafos pode capturar diferenças semânticas relevantes e fornece sinais complementares além da sobreposição lexical superficial para avaliar a qualidade da simplificação.
The use of textbooks as primary sources of information has increasingly given way to tools based on Large Language Models (LLMs), raising concerns about the reliability of generated answers. This study investigates how different adaptation strategies shape the behavior of small language models in educational Question Answering (QA) tasks in Portuguese. To support this analysis, we built a question-answer dataset derived from an NLP textbook and compared base models and the Retrieval-Augmented Generation (RAG) pipeline with models adapted through supervised fine-tuning and Continued Pretraining. The evaluation relies on questions from the LARI dataset, which has been validated by human specialists, and combines automatic and qualitative assessment procedures. The findings indicate that small models tuned with structured instructional knowledge achieve stronger semantic alignment and produce more pertinent answers in Portuguese educational QA scenarios.