Cecilia Bolaños
Author directory2026
ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
Šimon Sedláček | Sara Barahona | Cecilia Bolaños | Laura Herrera-Alarcón | Sathvik Udupa | Fernando López | Allison Ferner | Bolaji Yusuf | Alicia Lozano-Diez | Santosh Kesiraju | Ramani Duraiswami | Jan Černocký
Transactions of the Association for Computational Linguistics, Volume 14
Šimon Sedláček | Sara Barahona | Cecilia Bolaños | Laura Herrera-Alarcón | Sathvik Udupa | Fernando López | Allison Ferner | Bolaji Yusuf | Alicia Lozano-Diez | Santosh Kesiraju | Ramani Duraiswami | Jan Černocký
Transactions of the Association for Computational Linguistics, Volume 14
Reliable assessment of the abilities of large audio language models (LALMs) is essential to advancing the state of the art. As benchmarks rapidly evolve to incorporate complex reasoning and subjective tasks, they increasingly necessitate open-ended responses from LALMs. We present Open-ended Response Correctness Assessment (ORCA)—a reliable and lightweight model-based approach for answer correctness and disagreement modeling. We employ a three-stage annotation pipeline combining human judgment, structured feedback, and human-AI correction, yielding 9,663 annotations across 3,699 question-answer pairs from 15 LALMs on three audio understanding and reasoning benchmarks (achieving a Krippendorff’s alpha of 0.82). Our experiments employing curriculum learning show that ORCA models achieve a Spearman correlation of 0.91 with average human correctness ratings on seen benchmarks and generalize to unseen benchmarks with a score of 0.85, outperforming several LLM judge baselines including Gemini 2.5 Flash. Furthermore, we demonstrate that ORCA’s predicted variance correlates strongly with human disagreement, allowing it to effectively identify problematic benchmark items.
2025
Are Optimal Algorithms Still Optimal? Rethinking Sorting in LLM-Based Pairwise Ranking with Batching and Caching
Juan Wisznia | Cecilia Bolaños | Juan Tollo | Giovanni Franco Gabriel Marraffini | Agustín Andrés Gianolini | Noe Fabian Hsueh | Luciano Del Corro
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)
Juan Wisznia | Cecilia Bolaños | Juan Tollo | Giovanni Franco Gabriel Marraffini | Agustín Andrés Gianolini | Noe Fabian Hsueh | Luciano Del Corro
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)
We introduce a novel framework for analyzing sorting algorithms in pairwise ranking prompting (PRP), re-centering the cost model around LLM inferences rather than traditional pairwise comparisons. While classical metrics based on comparison counts have traditionally been used to gauge efficiency, our analysis reveals that expensive LLM inferences overturn these predictions; accordingly, our framework encourages strategies such as batching and caching to mitigate inference costs. We show that algorithms optimal in the classical setting can lose efficiency when LLM inferences dominate the cost under certain optimizations.