HuNeBR: A Multitask Benchmark to Evaluate LLMs’ Understanding of Northeastern Brazilian Portuguese Humor

José Matheus do Nascimento Gama, David Candeia Maia, Leandro Balby Marinho, Fabio Morais, João Brunet


Abstract
Humor recognition is a major challenge in Natural Language Processing (NLP) due to its subtle and context-dependent nature. Despite advances, Large Language Models (LLMs) still struggle with this task, especially in Brazilian Portuguese, where no dedicated benchmarks exist. This paper presents HuNeBR, a new benchmark of 475 annotated humorous texts from Northeastern Brazilian comedians. The benchmark evaluates LLMs on three tasks: identifying punchlines, classifying texts into eight comic styles, and explaining humor. This is the first benchmark to evaluate LLMs on the in-depth interpretation of humorous texts in Brazilian Portuguese, going beyond the binary tasks of traditional humor benchmarks. Both general-purpose and Portuguese-specialized LLMs were evaluated under zero-shot and few-shot settings. The findings indicate that LLMs perform very well at identifying punchlines, show inconsistent results in classifying comic styles, and produce humor interpretations that mostly align with human judgments. Among the models assessed, general-purpose multilingual systems like GPT-4 and Gemini 2.5 Flash achieved the top overall performance, whereas Sabiá 3.1, a model specialized in Brazilian Portuguese, demonstrated competitive results across all three tasks, highlighting the value of locally trained models in capturing linguistic and cultural subtleties.
Anthology ID:
2026.sigul-1.30
Volume:
Proceedings of the SIGUL 2026 Joint Workshop with ELE, EURALI, and DCLRL: Towards Inclusivity and Equality: Language Resources and Technologies for Under-Resourced and Endangered Languages
Month:
May
Year:
2026
Address:
Palma, Mallorca, Spain
Editors:
Atul Kr. Ojha, Sakriani Sakti, Claudia Soria, Maite Melero, John P. McCrae, Constantine Lignos, Chao-Hong Liu, German Rigau Claramunt, Georg Rehm
Venues:
SIGUL | EURALI | DCLRL | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
299–311
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-sigul-30
DOI:
10.63317/22b3hhj77udn
Bibkey:
Cite (ACL):
José Matheus do Nascimento Gama, David Candeia Maia, Leandro Balby Marinho, Fabio Morais, and João Brunet. 2026. HuNeBR: A Multitask Benchmark to Evaluate LLMs’ Understanding of Northeastern Brazilian Portuguese Humor. In Proceedings of the SIGUL 2026 Joint Workshop with ELE, EURALI, and DCLRL: Towards Inclusivity and Equality: Language Resources and Technologies for Under-Resourced and Endangered Languages, pages 299–311, Palma, Mallorca, Spain. ELRA Language Resources Association (ELRA).
Cite (Informal):
HuNeBR: A Multitask Benchmark to Evaluate LLMs’ Understanding of Northeastern Brazilian Portuguese Humor (Gama et al., SIGUL-EURALI-DCLRL 2026)
Copy Citation: