ORCA: Open-ended Response Correctness Assessment for Audio Question Answering

Šimon Sedláček, Sara Barahona, Cecilia Bolaños, Laura Herrera-Alarcón, Sathvik Udupa, Fernando López, Allison Ferner, Bolaji Yusuf, Alicia Lozano-Diez, Santosh Kesiraju, Ramani Duraiswami, Jan Černocký


Abstract
Reliable assessment of the abilities of large audio language models (LALMs) is essential to advancing the state of the art. As benchmarks rapidly evolve to incorporate complex reasoning and subjective tasks, they increasingly necessitate open-ended responses from LALMs. We present Open-ended Response Correctness Assessment (ORCA)—a reliable and lightweight model-based approach for answer correctness and disagreement modeling. We employ a three-stage annotation pipeline combining human judgment, structured feedback, and human-AI correction, yielding 9,663 annotations across 3,699 question-answer pairs from 15 LALMs on three audio understanding and reasoning benchmarks (achieving a Krippendorff’s alpha of 0.82). Our experiments employing curriculum learning show that ORCA models achieve a Spearman correlation of 0.91 with average human correctness ratings on seen benchmarks and generalize to unseen benchmarks with a score of 0.85, outperforming several LLM judge baselines including Gemini 2.5 Flash. Furthermore, we demonstrate that ORCA’s predicted variance correlates strongly with human disagreement, allowing it to effectively identify problematic benchmark items.
Anthology ID:
2026.tacl-1.100
Volume:
Transactions of the Association for Computational Linguistics, Volume 14
Month:
Year:
2026
Address:
Cambridge, MA
Venue:
TACL
SIG:
Publisher:
MIT Press
Note:
Pages:
2213–2233
Language:
URL:
https://aclanthology.org/2026.tacl-1.100/
DOI:
10.1162/tacl.a.798
Bibkey:
Cite (ACL):
Šimon Sedláček, Sara Barahona, Cecilia Bolaños, Laura Herrera-Alarcón, Sathvik Udupa, Fernando López, Allison Ferner, Bolaji Yusuf, Alicia Lozano-Diez, Santosh Kesiraju, Ramani Duraiswami, and Jan Černocký. 2026. ORCA: Open-ended Response Correctness Assessment for Audio Question Answering. Transactions of the Association for Computational Linguistics, 14:2213–2233.
Cite (Informal):
ORCA: Open-ended Response Correctness Assessment for Audio Question Answering (Sedláček et al., TACL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.tacl-1.100.pdf