SPINACH: SPARQL-Based Information Navigation for Challenging Real-World Questions

Shicheng Liu, Sina Semnani, Harold Triedman, Jialiang Xu, Isaac Zhao, Monica Lam


Abstract
Large Language Models (LLMs) have led to significant improvements in the Knowledge Base Question Answering (KBQA) task. However, datasets used in KBQA studies do not capture the true complexity of KBQA tasks. They either have simple questions, use synthetically generated logical forms, or are based on small knowledge base (KB) schemas.We introduce the SPINACH dataset, an expert-annotated KBQA dataset collected from discussions on Wikidata’s “Request a Query” forum with 320 decontextualized question-SPARQL pairs. The complexity of these in-the-wild queries calls for a KBQA system that can dynamically explore large and often incomplete schemas and reason about them, as it is infeasible to create a comprehensive training dataset. We also introduce an in-context learning KBQA agent, also called SPINACH, that mimics how a human expert would write SPARQLs to handle challenging questions. SPINACH achieves a new state of the art on the QALD-7, QALD-9 Plus and QALD-10 datasets by 31.0%, 27.0%, and 10.0% in F1, respectively, and coming within 1.6% of the fine-tuned LLaMA SOTA model on WikiWebQuestions.On our new SPINACH dataset, the SPINACH agent outperforms all baselines, including the best GPT-4-based KBQA agent, by at least 38.1% in F1.
Anthology ID:
2024.findings-emnlp.938
Volume:
Findings of the Association for Computational Linguistics: EMNLP 2024
Month:
November
Year:
2024
Address:
Miami, Florida, USA
Editors:
Yaser Al-Onaizan, Mohit Bansal, Yun-Nung Chen
Venue:
Findings
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
15977–16001
Language:
URL:
https://aclanthology.org/2024.findings-emnlp.938
DOI:
Bibkey:
Cite (ACL):
Shicheng Liu, Sina Semnani, Harold Triedman, Jialiang Xu, Isaac Zhao, and Monica Lam. 2024. SPINACH: SPARQL-Based Information Navigation for Challenging Real-World Questions. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 15977–16001, Miami, Florida, USA. Association for Computational Linguistics.
Cite (Informal):
SPINACH: SPARQL-Based Information Navigation for Challenging Real-World Questions (Liu et al., Findings 2024)
Copy Citation:
PDF:
https://aclanthology.org/2024.findings-emnlp.938.pdf
Software:
 2024.findings-emnlp.938.software.zip
Data:
 2024.findings-emnlp.938.data.zip