The GELATO Dataset for Legislative NER

Matthew Flynn, Timothy Obiso, Sam Newman


Abstract
This paper introduces GELATO (Government, Executive, Legislative, and Treaty Ontology), a dataset of U.S. House and Senate bills from the 118th Congress annotated using a novel two-level named entity recognition ontology designed for U.S. legislative texts. We fine-tune transformer-based models (BERT, RoBERTa) of different architectures and sizes on this dataset for first-level prediction. We then use LLMs with optimized prompts to complete the second level prediction. The strong performance of RoBERTa and relatively weak performance of BERT models, as well as the application of LLMs as second-level predictors, support future research in legislative NER or downstream tasks using these model combinations as extraction tools.
Anthology ID:
2026.lrec-1.569
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
7163–7177
Language:
External URL:
https://lrec.elra.info/lrec2026-main-569
DOI:
10.63317/3axxkz9oh5th
Bibkey:
Cite (ACL):
Matthew Flynn, Timothy Obiso, and Sam Newman. 2026. The GELATO Dataset for Legislative NER. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 7163–7177, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
The GELATO Dataset for Legislative NER (Flynn et al., LREC 2026)
Copy Citation: