Normalizing Section Names and Structure of Scientific Articles

Nicolau Duran-Silva, Julian Moreno-Schneider, César A. Parra-Rojas, Georg Rehm


Abstract
The growing amount of scientific literature has increased the need for automatic methods that can retrieve, process, and exploit scholarly content. In this work, we explore section name normalization and hierarchy prediction for scientific articles using a two-level taxonomy. We compare independent, sequential classification models, and generative large language models on the SASC dataset. Results show that classification approaches, particularly sequential models that employ document-level context, consistently outperform generative methods. Incorporating section content is essential for fine-grained classification, while generative models remain limited in zero-shot settings. Our experiments highlight the importance of structure-aware modelling for large-scale scholarly document processing, and the importance of section normalization for the development of advanced research mapping and research assessment tools.
Anthology ID:
2026.nslp-1.21
Volume:
Proceedings of Natural Scientific Language Processing (NSLP) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Georg Rehm, Stefan Dietze, Danilo Dessi, Diana Maynard, Sonja Schimmler
Venues:
NSLP | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
218–224
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-nslp-21
DOI:
10.63317/5ftk2fxxf7jd
Bibkey:
Cite (ACL):
Nicolau Duran-Silva, Julian Moreno-Schneider, César A. Parra-Rojas, and Georg Rehm. 2026. Normalizing Section Names and Structure of Scientific Articles. In Proceedings of Natural Scientific Language Processing (NSLP) @ LREC 2026, pages 218–224, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
Normalizing Section Names and Structure of Scientific Articles (Duran-Silva et al., NSLP 2026)
Copy Citation: