Is This Idea Novel? An Automated Benchmark for Judgment of Research Ideas

Tim Schopf, Michael Färber


Abstract
Judging the novelty of research ideas is crucial for advancing science, enabling the identification of unexplored directions, and ensuring contributions meaningfully extend existing knowledge rather than reiterate minor variations. However, given the exponential growth of scientific literature, manually judging the novelty of research ideas through literature reviews is labor-intensive, subjective, and infeasible at scale. Therefore, recent efforts have proposed automated approaches for research idea novelty judgment. Yet, evaluation of these approaches remains largely inconsistent and is typically based on non-standardized human evaluations, hindering large-scale, comparable evaluations. To address this, we introduce RINoBench, the first comprehensive benchmark for large-scale evaluation of research idea novelty judgments. It comprises 1,381 research ideas derived from and judged by human experts as well as nine automated evaluation metrics designed to assess both rubric-based novelty scores and textual justifications of novelty judgments. Using this benchmark, we evaluate several state-of-the-art large language models (LLMs) on their ability to judge the novelty of research ideas. Our findings reveal that while LLM-generated reasoning closely mirrors human rationales, this alignment does not reliably translate into accurate novelty judgments, which diverge significantly from human gold standard judgments—even among leading reasoning-capable models. Data and code available at: https://github.com/TimSchopf/RINoBench
Anthology ID:
2026.lrec-1.370
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
4716–4727
Language:
External URL:
https://lrec.elra.info/lrec2026-main-370
DOI:
10.63317/4c3gy3f7epnj
Bibkey:
Cite (ACL):
Tim Schopf and Michael Färber. 2026. Is This Idea Novel? An Automated Benchmark for Judgment of Research Ideas. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 4716–4727, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Is This Idea Novel? An Automated Benchmark for Judgment of Research Ideas (Schopf & Färber, LREC 2026)
Copy Citation:
Optionalsupplementarymaterial:
 2026.lrec-1.370.OptionalSupplementaryMaterial.zip