Reformulating Information Retrieval from Speech and Text as a Detection Problem

Damianos Karakos, Rabih Zbib, William Hartmann, Richard Schwartz, John Makhoul


Abstract
In the IARPA MATERIAL program, information retrieval (IR) is treated as a hard detection problem; the system has to output a single global ranking over all queries, and apply a hard threshold on this global list to come up with all the hypothesized relevant documents. This means that how queries are ranked relative to each other can have a dramatic impact on performance. In this paper, we study such a performance measure, the Average Query Weighted Value (AQWV), which is a combination of miss and false alarm rates. AQWV requires that the same detection threshold is applied to all queries. Hence, detection scores of different queries should be comparable, and, to do that, a score normalization technique (commonly used in keyword spotting from speech) should be used. We describe unsupervised methods for score normalization, which are borrowed from the speech field and adapted accordingly for IR, and demonstrate that they greatly improve AQWV on the task of cross-language information retrieval (CLIR), on three low-resource languages used in MATERIAL. We also present a novel supervised score normalization approach which gives additional gains.
Anthology ID:
2020.clssts-1.7
Volume:
Proceedings of the workshop on Cross-Language Search and Summarization of Text and Speech (CLSSTS2020)
Month:
May
Year:
2020
Address:
Marseille, France
Editors:
Kathy McKeown, Douglas W. Oard, Elizabeth, Richard Schwartz
Venue:
CLSSTS
SIG:
Publisher:
European Language Resources Association
Note:
Pages:
38–43
Language:
English
URL:
https://aclanthology.org/2020.clssts-1.7
DOI:
Bibkey:
Cite (ACL):
Damianos Karakos, Rabih Zbib, William Hartmann, Richard Schwartz, and John Makhoul. 2020. Reformulating Information Retrieval from Speech and Text as a Detection Problem. In Proceedings of the workshop on Cross-Language Search and Summarization of Text and Speech (CLSSTS2020), pages 38–43, Marseille, France. European Language Resources Association.
Cite (Informal):
Reformulating Information Retrieval from Speech and Text as a Detection Problem (Karakos et al., CLSSTS 2020)
Copy Citation:
PDF:
https://aclanthology.org/2020.clssts-1.7.pdf