Struct2Unstruct: Creating Tender NER Datasets from Structured Procurement Records using Large Language Models

Asim Abbas, Mark Lee, Niloofer Shanavas, Venelin Kovatchev, Mubashir Ali


Abstract
Named Entity Recognition (NER) in the tender and procurement domain is critical for tasks such as contract monitoring, supplier analysis, and compliance tracking. However, unlike general-purpose NER, no open-source datasets exist for Tender NER, largely due to data sensitivity and confidentiality restrictions. This scarcity limits the development of automated entity extraction models. To address this gap, we propose struct2unstruct, a data preparation pipeline that generates and annotates tender-specific datasets using large language models (LLMs). Starting from structured procurement data published by the Singapore government (2015–2021) available in English language, we employ Llama-3 to generate synthetic tender narratives in multiple writing styles, ensuring each contains at least one tender-related entity. Post-processing steps correct inconsistencies in dates, symbols, and entity formats. Entities are then annotated using a BIO tagging scheme through deterministic alignment with structured fields, followed by expert validation to ensure accuracy. This study focuses on data preparation and evaluation, not model training. The resulting dataset provides a scalable resource for future Tender NER research in low-resource environments. By releasing both the dataset and pipeline as open-source resources, we establish a foundation for advancing domain-adapted information extraction and automated tender entity recognition.
Anthology ID:
2026.resourceful-4.10
Volume:
Proceedings of the Fourth Workshop on the Role of Resources in the Age of Large Language Models (RESOURCEFUL 2026)
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Felix Morger, Nikolai Ilinykh, Barbara Scalvini, Simon Dobnik, Dana Dannélls
Venues:
RESOURCEFUL | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
96–106
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-resourceful-10
DOI:
10.63317/3mtwxjwhaqus
Bibkey:
Cite (ACL):
Asim Abbas, Mark Lee, Niloofer Shanavas, Venelin Kovatchev, and Mubashir Ali. 2026. Struct2Unstruct: Creating Tender NER Datasets from Structured Procurement Records using Large Language Models. In Proceedings of the Fourth Workshop on the Role of Resources in the Age of Large Language Models (RESOURCEFUL 2026), pages 96–106, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
Struct2Unstruct: Creating Tender NER Datasets from Structured Procurement Records using Large Language Models (Abbas et al., RESOURCEFUL 2026)
Copy Citation: