DeepQuestion: Systematic Generation of Real-World Challenges for Evaluating LLMs Performance

Ali Khoramfar, Ali Ramezani, Mohammad Mahdi Mohajeri, Mohammad Javad Dousti, Majid Nili Ahmadabadi, Heshaam Faili


Abstract
While Large Language Models (LLMs) achieve near-human performance on standard benchmarks, their capabilities often fail to generalize to complex, real-world problems. To bridge this gap, we introduce DeepQuestion, a scalable, automated framework that systematically elevates the cognitive complexity of existing datasets through controlled task transformations grounded in explicit cognitive hierarchies. Based on Bloom’s taxonomy, DeepQuestion generates (1) scenario-based problems to test the application of knowledge in noisy, realistic contexts, and (2) instruction-based prompts that require models to create new questions from a given solution path, assessing synthesis and evaluation skills. Our extensive evaluation across ten leading open-source and proprietary models, covering both general-purpose and reasoning LLMs, reveals a stark performance decline—with accuracy dropping by up to 70%—as tasks ascend the cognitive hierarchy across evaluation settings. These findings underscore that current benchmarks overestimate true reasoning abilities and highlight the critical need for cognitively diverse evaluations to guide future LLM development.
Anthology ID:
2026.lrec-1.896
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
11451–11460
Language:
External URL:
https://lrec.elra.info/lrec2026-main-896
DOI:
10.63317/37w7pv6oaaeg
Bibkey:
Cite (ACL):
Ali Khoramfar, Ali Ramezani, Mohammad Mahdi Mohajeri, Mohammad Javad Dousti, Majid Nili Ahmadabadi, and Heshaam Faili. 2026. DeepQuestion: Systematic Generation of Real-World Challenges for Evaluating LLMs Performance. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 11451–11460, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
DeepQuestion: Systematic Generation of Real-World Challenges for Evaluating LLMs Performance (Khoramfar et al., LREC 2026)
Copy Citation:
Optionalsupplementarymaterial:
 2026.lrec-1.896.OptionalSupplementaryMaterial.zip