Benchmarking LLMs for ARR Area Assignment: Evidence and Implications for Assignment Strategies

Eileen Bingert, Diego Alves, Stefania Degaetano-Ortlieb


Abstract
We study how large language models (LLMs) perform at assigning ACL Rolling Review (ARR) areas from paper titles/abstracts. Using 558 papers (ACL/EACL/NAACL, 2020 to 2025), we compare multiple LLMs and prompting schemes (zero/few-shot; with/without ARR keywords; each-category variants) and analyze per-area scores, error overlap, and confusion matrices. One-shot prompting (with OpenAI-gpt-oss-20b) tends to perform best, while injecting ARR keywords often lowers accuracy. Task-bounded areas (e.g., MT, IE, QA, Summarization) are predicted more reliably, whereas broad, cross-cutting labels (e.g., Resources and Evaluation, NLP Applications) are frequently conflated, indicating taxonomy ambiguity rather than solely model limitations. We recommend hierarchical or primary-plus-secondary labels to reduce ambiguity and improve reviewer matching. Our dataset, methods, and findings offer a reproducible baseline for area selection support in ACL workflows.
Anthology ID:
2026.nslp-1.2
Volume:
Proceedings of Natural Scientific Language Processing (NSLP) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Georg Rehm, Stefan Dietze, Danilo Dessi, Diana Maynard, Sonja Schimmler
Venues:
NSLP | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
13–24
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-nslp-02
DOI:
10.63317/4jbm8o3us4c8
Bibkey:
Cite (ACL):
Eileen Bingert, Diego Alves, and Stefania Degaetano-Ortlieb. 2026. Benchmarking LLMs for ARR Area Assignment: Evidence and Implications for Assignment Strategies. In Proceedings of Natural Scientific Language Processing (NSLP) @ LREC 2026, pages 13–24, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
Benchmarking LLMs for ARR Area Assignment: Evidence and Implications for Assignment Strategies (Bingert et al., NSLP 2026)
Copy Citation: