Empirical Comparisons of MASC Word Sense Annotations

Gerard de Melo, Collin F. Baker, Nancy Ide, Rebecca J. Passonneau, Christiane Fellbaum


Abstract
We analyze how different conceptions of lexical semantics affect sense annotations and how multiple sense inventories can be compared empirically, based on annotated text. Our study focuses on the MASC project, where data has been annotated using WordNet sense identifiers on the one hand, and FrameNet lexical units on the other. This allows us to compare the sense inventories of these lexical resources empirically rather than just theoretically, based on their glosses, leading to new insights. In particular, we compute contingency matrices and develop a novel measure, the Expected Jaccard Index, that quantifies the agreement between annotations of the same data based on two different resources even when they have different sets of categories.
Anthology ID:
L12-1525
Volume:
Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
Month:
May
Year:
2012
Address:
Istanbul, Turkey
Editors:
Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
3036–3043
Language:
URL:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/884_Paper.pdf
DOI:
Bibkey:
Cite (ACL):
Gerard de Melo, Collin F. Baker, Nancy Ide, Rebecca J. Passonneau, and Christiane Fellbaum. 2012. Empirical Comparisons of MASC Word Sense Annotations. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12), pages 3036–3043, Istanbul, Turkey. European Language Resources Association (ELRA).
Cite (Informal):
Empirical Comparisons of MASC Word Sense Annotations (de Melo et al., LREC 2012)
Copy Citation:
PDF:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/884_Paper.pdf