Structured Tender Entities Extraction from Complex Tables with Few-short Learning

Asim Abbas; Mark Lee; Niloofer Shanavas; Venelin Kovatchev; Mubashir Ali

Structured Tender Entities Extraction from Complex Tables with Few-short Learning

Asim Abbas, Mark Lee, Niloofer Shanavas, Venelin Kovatchev, Mubashir Ali

Abstract

Extracting structured text from complex tables in PDF tender documents remains a challenging task due to the loss of structural and positional information during the extraction process. AI-based models often require extensive training data, making development from scratch both tedious and time-consuming. Our research focuses on identifying tender entities in complex table formats within PDF documents. To address this, we propose a novel approach utilizing few-shot learning with large language models (LLMs) to restore the structure of extracted text. Additionally, handcrafted rules and regular expressions are employed for precise entity classification. To evaluate the robustness of LLMs with few-shot learning, we employ data-shuffling techniques. Our experiments show that current text extraction tools fail to deliver satisfactory results for complex table structures. However, the few-shot learning approach significantly enhances the structural integrity of extracted data and improves the accuracy of tender entity identification.

Anthology ID:: 2025.regnlp-1.9
Volume:: Proceedings of the 1st Regulatory NLP Workshop (RegNLP 2025)
Month:: January
Year:: 2025
Address:: Abu Dhabi, UAE
Editors:: Tuba Gokhan, Kexin Wang, Iryna Gurevych, Ted Briscoe
Venues:: RegNLP | WS
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 59–67
Language:
URL:: https://aclanthology.org/2025.regnlp-1.9/
DOI:
Bibkey:
Cite (ACL):: Asim Abbas, Mark Lee, Niloofer Shanavas, Venelin Kovatchev, and Mubashir Ali. 2025. Structured Tender Entities Extraction from Complex Tables with Few-short Learning. In Proceedings of the 1st Regulatory NLP Workshop (RegNLP 2025), pages 59–67, Abu Dhabi, UAE. Association for Computational Linguistics.
Cite (Informal):: Structured Tender Entities Extraction from Complex Tables with Few-short Learning (Abbas et al., RegNLP 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.regnlp-1.9.pdf

PDF Cite Search Fix data