From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation

Aviya Maimon, Amir David Nisan Cohen, Gal Vishne, Shauli Ravfogel, Reut Tsarfaty


Abstract
Current evaluations of large language models (LLMs) rely heavily on a growing collection of benchmarks and on aggregate benchmark scores, yet it remains unclear what this comparison actually captures, and what these scores reveal about models’ underlying capabilities. Here, we propose a new paradigm for LLM evaluation, by asking whether benchmark performance reflects many independent abilities, or rather, relies on a small number of shared dimensions. To answer this, we apply Factor Analysis (FA) to a massive performance matrix of LLMs versus benchmarks (60 × 44) revealing an intrinsically low-rank structure of that matrix. That is, a small number of latent factors captures most of the structure in the full task space. This low-rank geometry reveals substantial redundancy across existing tasks and explains why many benchmarks appear to be measuring overlapping abilities. We further show that these latent factors correspond to coherent, skill-like, dimensions of LLM behavior. Leveraging this latent skill-space, we deliver three practical tools for LLM evaluation and downstream users: (i) identifying redundant tasks, (ii) profiling new models using a small subset of tasks, and (iii) selecting models aligned with desired skill profiles. Our method provides a solid alternative to the de-facto standard of a single aggregate score, and establishes an interpretable and practical framework for understanding and benchmarking LLM core capabilities.
Anthology ID:
2026.tacl-1.75
Volume:
Transactions of the Association for Computational Linguistics, Volume 14
Month:
Year:
2026
Address:
Cambridge, MA
Venue:
TACL
SIG:
Publisher:
MIT Press
Note:
Pages:
1660–1684
Language:
URL:
https://aclanthology.org/2026.tacl-1.75/
DOI:
10.1162/tacl.a.745
Bibkey:
Cite (ACL):
Aviya Maimon, Amir David Nisan Cohen, Gal Vishne, Shauli Ravfogel, and Reut Tsarfaty. 2026. From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation. Transactions of the Association for Computational Linguistics, 14:1660–1684.
Cite (Informal):
From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation (Maimon et al., TACL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.tacl-1.75.pdf