A Taxonomy of Safety: Harmonizing LLM Benchmarks in a Fragmented Landscape

Shadi Rastegar, Viktor Hangya, Fabian Kuech, Darina Gold


Abstract
Understanding and mitigating the safety limitations of LLMs is of great importance to build trustworthy AI applications. Although a wide range of safety benchmarks are available, there is no standardized taxonomy of safety categories. As a result, some benchmarks focus on a specific subset of categories, they define test samples on different granularity levels, or they use different definitions or naming conventions. To mitigate these issues, we propose a two-level taxonomy of LLM safety categories, created by harmonizing existing resources. Our taxonomy gives an overview of important safety categories that helps researchers pinpoint potential safety risks and select the right benchmarks when evaluating or developing language models. Moreover, the taxonomy provides guidelines to categorize future benchmarks. Furthermore, since the majority of the available safety resources are English-focused, we check the cross-cultural validity of our taxonomy by translating datasets covering all top level categories to French, German, Italian, and Spanish. A manual review of a subset of translated samples by native speakers revealed no major cultural mismatches from a safety perspective. This supports not only the transferability of English benchmarks but also the transferability of the categories in our taxonomy, as well as its potential as a practical tool for guiding safety-focused dataset development and evaluation beyond English.
Anthology ID:
2026.lrec-1.350
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
4470–4481
Language:
External URL:
https://lrec.elra.info/lrec2026-main-350
DOI:
10.63317/4n7jrxunmvcp
Bibkey:
Cite (ACL):
Shadi Rastegar, Viktor Hangya, Fabian Kuech, and Darina Gold. 2026. A Taxonomy of Safety: Harmonizing LLM Benchmarks in a Fragmented Landscape. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 4470–4481, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
A Taxonomy of Safety: Harmonizing LLM Benchmarks in a Fragmented Landscape (Rastegar et al., LREC 2026)
Copy Citation: