Do LLM Judges Agree with Humans? Evaluating Financial Commentaries from Material Facts

Gabriel Assis, Daniela Vianna, Livy Real, Marina Ramalhete Masid, João Nepomuceno, Eduardo Rottschaefer, Altigran Soares da Silva, Aline Paes


Abstract
The rapid adoption of large language models (LLMs) in finance has enabled automated generation of in-domain texts such as earnings summaries, market analyses, and commentary on regulated disclosures. Generating accurate and accessible financial commentary from material facts poses challenges related to domain-specific language, strict factual faithfulness, and readability. Evaluating such outputs is difficult: traditional automatic metrics overlook financial correctness and adequacy, while human expert assessment is reliable but costly and subjective. This paper proposes a multidimensional evaluation protocol for generating financial commentary in Portuguese and investigates the use of LLMs as evaluators (“LLMs-as-judges”) in this high-stakes, low-resource setting. We systematically compare human expert judgments and LLM-based evaluation preferences, while also investigating complementary dimensions such as writing quality, factuality, usefulness, and simplicity. To the best of our knowledge, this is the first proposal of an evaluation protocol for this task in Portuguese. Furthermore, we analyze where LLM judgments align with or diverge from human assessments, providing practical insights and recommendations for evaluation methodologies in financial AI applications.
Anthology ID:
2026.stil-1.3
Volume:
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Month:
October
Year:
2026
Address:
Cuiabá, Mato Grosso, Brazil
Editors:
Bryan Khelven da Silva Barbosa, Aline Paes, Ariani Di Felippo
Venue:
STIL
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
26–40
Language:
URL:
https://aclanthology.org/2026.stil-1.3/
DOI:
10.5753/stil.2026.26484
Bibkey:
Cite (ACL):
Gabriel Assis, Daniela Vianna, Livy Real, Marina Ramalhete Masid, João Nepomuceno, Eduardo Rottschaefer, Altigran Soares da Silva, and Aline Paes. 2026. Do LLM Judges Agree with Humans? Evaluating Financial Commentaries from Material Facts. In Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology, pages 26–40, Cuiabá, Mato Grosso, Brazil. Association for Computational Linguistics.
Cite (Informal):
Do LLM Judges Agree with Humans? Evaluating Financial Commentaries from Material Facts (Assis et al., STIL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.stil-1.3.pdf