Identifying Nuggets of Information in GALE Distillation Evaluation

Olga Babko-Malaya, Greg Milette, Michael Schneider, Sarah Scogin


Abstract
This paper describes an approach to automatic nuggetization and implemented system employed in GALE Distillation evaluation to measure the information content of text returned in response to an open-ended question. The system identifies nuggets, or atomic units of information, categorizes them according to their semantic type, and selects different types of nuggets depending on the type of the question. We further show how this approach addresses the main challenges for using automatic nuggetization for QA evaluation: the variability of relevant nuggets and their dependence on the question. Specifically, we propose a template-based approach to nuggetization, where different semantic categories of nuggets are extracted dependent on the template of a question. During evaluation, human annotators judge each snippet returned in response to a query as relevant or irrelevant, whereas automatic template-based nuggetization is further used to identify the semantic units of information that people would have selected as ‘relevant' or ‘irrelevant' nuggets for a given query. Finally, the paper presents the performance results of the nuggetization system which compare the number of automatically generated nuggets and human nuggets and show that our automatic nuggetization is consistent with human judgments.
Anthology ID:
L12-1261
Volume:
Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
Month:
May
Year:
2012
Address:
Istanbul, Turkey
Editors:
Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
2322–2327
Language:
URL:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/482_Paper.pdf
DOI:
Bibkey:
Cite (ACL):
Olga Babko-Malaya, Greg Milette, Michael Schneider, and Sarah Scogin. 2012. Identifying Nuggets of Information in GALE Distillation Evaluation. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12), pages 2322–2327, Istanbul, Turkey. European Language Resources Association (ELRA).
Cite (Informal):
Identifying Nuggets of Information in GALE Distillation Evaluation (Babko-Malaya et al., LREC 2012)
Copy Citation:
PDF:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/482_Paper.pdf