João Nepomuceno
Author directory2026
Do LLM Judges Agree with Humans? Evaluating Financial Commentaries from Material Facts
Gabriel Assis | Daniela Vianna | Livy Real | Marina Ramalhete Masid | João Nepomuceno | Eduardo Rottschaefer | Altigran Soares da Silva | Aline Paes
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Gabriel Assis | Daniela Vianna | Livy Real | Marina Ramalhete Masid | João Nepomuceno | Eduardo Rottschaefer | Altigran Soares da Silva | Aline Paes
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
The rapid adoption of large language models (LLMs) in finance has enabled automated generation of in-domain texts such as earnings summaries, market analyses, and commentary on regulated disclosures. Generating accurate and accessible financial commentary from material facts poses challenges related to domain-specific language, strict factual faithfulness, and readability. Evaluating such outputs is difficult: traditional automatic metrics overlook financial correctness and adequacy, while human expert assessment is reliable but costly and subjective. This paper proposes a multidimensional evaluation protocol for generating financial commentary in Portuguese and investigates the use of LLMs as evaluators (“LLMs-as-judges”) in this high-stakes, low-resource setting. We systematically compare human expert judgments and LLM-based evaluation preferences, while also investigating complementary dimensions such as writing quality, factuality, usefulness, and simplicity. To the best of our knowledge, this is the first proposal of an evaluation protocol for this task in Portuguese. Furthermore, we analyze where LLM judgments align with or diverge from human assessments, providing practical insights and recommendations for evaluation methodologies in financial AI applications.