PAUQ: Text-to-SQL in Russian

Daria Bakshandaeva, Oleg Somov, Ekaterina Dmitrieva, Vera Davydova, Elena Tutubalina


Abstract
Semantic parsing is an important task that allows to democratize human-computer interaction. One of the most popular text-to-SQL datasets with complex and diverse natural language (NL) questions and SQL queries is Spider. We construct and complement a Spider dataset for Russian, thus creating the first publicly available text-to-SQL dataset for this language. While examining its components - NL questions, SQL queries and databases content - we identify limitations of the existing database structure, fill out missing values for tables and add new requests for underrepresented categories. We select thirty functional test sets with different features that can be used for the evaluation of neural models’ abilities. To conduct the experiments, we adapt baseline architectures RAT-SQL and BRIDGE and provide in-depth query component analysis. On the target language, both models demonstrate strong results with monolingual training and improved accuracy in multilingual scenario. In this paper, we also study trade-offs between machine-translated and manually-created NL queries. At present, Russian text-to-SQL is lacking in datasets as well as trained models, and we view this work as an important step towards filling this gap.
Anthology ID:
2022.findings-emnlp.175
Volume:
Findings of the Association for Computational Linguistics: EMNLP 2022
Month:
December
Year:
2022
Address:
Abu Dhabi, United Arab Emirates
Editors:
Yoav Goldberg, Zornitsa Kozareva, Yue Zhang
Venue:
Findings
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
2355–2376
Language:
URL:
https://aclanthology.org/2022.findings-emnlp.175
DOI:
10.18653/v1/2022.findings-emnlp.175
Bibkey:
Cite (ACL):
Daria Bakshandaeva, Oleg Somov, Ekaterina Dmitrieva, Vera Davydova, and Elena Tutubalina. 2022. PAUQ: Text-to-SQL in Russian. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 2355–2376, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
Cite (Informal):
PAUQ: Text-to-SQL in Russian (Bakshandaeva et al., Findings 2022)
Copy Citation:
PDF:
https://aclanthology.org/2022.findings-emnlp.175.pdf