Flavio Carvalho

Author directory

2026

This paper evaluates automated feedback generation for ENEM Competency 5 by augmenting zero-shot prompts with an Evaluator Term Set (CTA), a term set proposed in this work and extracted from human evaluator comments on Competency 5. We compare CTA-augmented feedback against a zero-shot baseline on 46 essays with human reference feedback using BERTScore F1 and the Wilcoxon signed-rank test. CTA augmentation increases mean BERTScore F1 for both models, with statistically significant improvement in semantic similarity to human feedback for both evaluated models at α = 0.05.