PolyglotQL: A Pipeline for Multilingual Text-to-SPARQL Dataset Generation

Julio Perez, Fabio Barth, Georg Rehm


Abstract
We present PolyglotQL, an open-source ETL (Extract, Transform, Load) pipeline for systematically creating multilingual text-to-SPARQL datasets, along with an accompanying framework for evaluating text-to-SPARQL generation models. PolyglotQL provides an extensible and modular architecture that aggregates, normalizes, and augments heterogeneous question–SPARQL pairs from established text-to-SPARQL datasets. With this pipeline, we automatically construct a bilingual English–German dataset featuring contextualized entity and relationship mappings as well as automatically translated and aligned question pairs. We also conduct an empirical evaluation using two multilingual open large language models under two distinct contextualization settings. The results show consistent performance improvements when explicit grounding information is provided, highlighting the benefits of structured context in multilingual semantic parsing.
Anthology ID:
2026.lrec-1.531
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
6674–6684
Language:
External URL:
https://lrec.elra.info/lrec2026-main-531
DOI:
10.63317/5ow3k3fbz296
Bibkey:
Cite (ACL):
Julio Perez, Fabio Barth, and Georg Rehm. 2026. PolyglotQL: A Pipeline for Multilingual Text-to-SPARQL Dataset Generation. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 6674–6684, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
PolyglotQL: A Pipeline for Multilingual Text-to-SPARQL Dataset Generation (Perez et al., LREC 2026)
Copy Citation: