Backtracing: Retrieving the Cause of the Query

Rose Wang; Pawan Wirawarn; Omar Khattab; Noah Goodman; Dorottya Demszky

Backtracing: Retrieving the Cause of the Query

Rose Wang, Pawan Wirawarn, Omar Khattab, Noah Goodman, Dorottya Demszky

Abstract

Many online content portals allow users to ask questions to supplement their understanding (e.g., of lectures). While information retrieval (IR) systems may provide answers for such user queries, they do not directly assist content creators—such as lecturers who want to improve their content—identify segments that caused a user to ask those questions.We introduce the task of backtracing, in which systems retrieve the text segment that most likely caused a user query.We formalize three real-world domains for which backtracing is important in improving content delivery and communication: understanding the cause of (a) student confusion in the Lecture domain, (b) reader curiosity in the News Article domain, and (c) user emotion in the Conversation domain.We evaluate the zero-shot performance of popular information retrieval methods and language modeling methods, including bi-encoder, re-ranking and likelihood-based methods and ChatGPT.While traditional IR systems retrieve semantically relevant information (e.g., details on “projection matrices” for a query “does projecting multiple times still lead to the same point?”), they often miss the causally relevant context (e.g., the lecturer states “projecting twice gets me the same answer as one projection”). Our results show that there is room for improvement on backtracing and it requires new retrieval approaches.We hope our benchmark serves to improve future retrieval systems for backtracing, spawning systems that refine content generation and identify linguistic triggers influencing user queries.

Anthology ID:: 2024.findings-eacl.48
Volume:: Findings of the Association for Computational Linguistics: EACL 2024
Month:: March
Year:: 2024
Address:: St. Julian’s, Malta
Editors:: Yvette Graham, Matthew Purver
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 722–735
Language:
URL:: https://aclanthology.org/2024.findings-eacl.48
DOI:
Bibkey:
Cite (ACL):: Rose Wang, Pawan Wirawarn, Omar Khattab, Noah Goodman, and Dorottya Demszky. 2024. Backtracing: Retrieving the Cause of the Query. In Findings of the Association for Computational Linguistics: EACL 2024, pages 722–735, St. Julian’s, Malta. Association for Computational Linguistics.
Cite (Informal):: Backtracing: Retrieving the Cause of the Query (Wang et al., Findings 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.findings-eacl.48.pdf
Video:: https://aclanthology.org/2024.findings-eacl.48.mp4

PDF Cite Search Video