Using LLMs for Automatic Discipline Annotation in a Diachronic Corpus of English Scientific Papers

Sergei Bagdasarov, Diego Alves, Stefan Fischer, Elke Teich


Abstract
This study investigates the potential of generative large language models (LLMs) to automatically identify the disciplines of scientific papers in the Royal Society Corpus (RSC) – an extensive collection of English scientific publications spanning more than three centuries. We evaluated eight open-source, state-of-the-art LLMs from four model families on a manually annotated subset and further validated the three best-performing models on a corpus of modern scientific texts. These models were subsequently used for large-scale annotation of the RSC. The models exhibited robust and consistent performance, with at least two LLMs agreeing on the same label for 98.3% of the documents. We then conducted an error analysis of papers assigned divergent labels and a diachronic case study of disciplinary trends within the corpus. The error analysis revealed that most discrepancies occurred in twentieth-century texts, reflecting the growing interdisciplinarity of research. The diachronic analysis showed a gradual decline in disciplinary diversity over time as well as fluctuations corresponding to major paradigm shifts such as the Chemical Revolution and key twentieth-century developments in Physics. The discipline labels generated by the three models will be made publicly available.
Anthology ID:
2026.lrec-1.187
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
2376–2386
Language:
External URL:
https://lrec.elra.info/lrec2026-main-187
DOI:
10.63317/3j9wvu86v48t
Bibkey:
Cite (ACL):
Sergei Bagdasarov, Diego Alves, Stefan Fischer, and Elke Teich. 2026. Using LLMs for Automatic Discipline Annotation in a Diachronic Corpus of English Scientific Papers. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 2376–2386, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Using LLMs for Automatic Discipline Annotation in a Diachronic Corpus of English Scientific Papers (Bagdasarov et al., LREC 2026)
Copy Citation: