LLM Evaluation in Practice: A Review of Metrics, Practitioner Insights, and Lessons Learned

Roos M. Bakker, Marianne Witte-Schaaphok, Julia García-Fernández, Tom Brand, Jens van der Weide, Stephan Raaijmakers


Abstract
The rapid, widespread adoption of Large Language Models (LLMs) highlights the need to understand their performance, strengths, and limitations. However, evaluating LLMs presents significant challenges due to the broad range of tasks and model capabilities, especially in practice or low-resource settings where benchmark datasets are not available. In text generation tasks, answer diversity has always complicated automatic evaluation, and the enhanced fluency and creativity of LLMs lead to further challenges. Existing metrics and frameworks often fail to account for these complexities. Furthermore, recent research into the replicability of benchmarks has demonstrated serious issues when reproducing historical benchmark results. This paper makes two key contributions: (1) a categorisation of challenges and metrics in LLM evaluation, and (2) lessons learned from practice through a survey and a use case. To this end, a literature study was conducted to identify challenges and metrics in scientific work. A survey among developers working with LLMs provided insights into practical challenges. Furthermore, selected metrics were implemented in a practical use case to gain insights into their strengths and limitations. By combining theoretical analysis with real-world experiences and lessons learned from practice, this work provides an overview and best practices for users evaluating LLM performance.
Anthology ID:
2026.llms4ssh-1.3
Volume:
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma de Mallorca (Spain)
Editors:
Arturo Montejo-Raez, Cristina Grisot, Joanna Blochowiak, Nikola Ljubešić, Elena Battaner, German Rigau
Venues:
LLMs4SSH | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
23–38
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-llms4ssh-03
DOI:
10.63317/3exz7ndz9d6j
Bibkey:
Cite (ACL):
Roos M. Bakker, Marianne Witte-Schaaphok, Julia García-Fernández, Tom Brand, Jens van der Weide, and Stephan Raaijmakers. 2026. LLM Evaluation in Practice: A Review of Metrics, Practitioner Insights, and Lessons Learned. In Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026, pages 23–38, Palma de Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
LLM Evaluation in Practice: A Review of Metrics, Practitioner Insights, and Lessons Learned (Bakker et al., LLMs4SSH 2026)
Copy Citation: