Piper Armstrong
2019
Cross-referencing Using Fine-grained Topic Modeling
Jeffrey Lund
|
Piper Armstrong
|
Wilson Fearn
|
Stephen Cowley
|
Emily Hales
|
Kevin Seppi
Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)
Cross-referencing, which links passages of text to other related passages, can be a valuable study aid for facilitating comprehension of a text. However, cross-referencing requires first, a comprehensive thematic knowledge of the entire corpus, and second, a focused search through the corpus specifically to find such useful connections. Due to this, cross-reference resources are prohibitively expensive and exist only for the most well-studied texts (e.g. religious texts). We develop a topic-based system for automatically producing candidate cross-references which can be easily verified by human annotators. Our system utilizes fine-grained topic modeling with thousands of highly nuanced and specific topics to identify verse pairs which are topically related. We demonstrate that our system can be cost effective compared to having annotators acquire the expertise necessary to produce cross-reference resources unaided.
Automatic Evaluation of Local Topic Quality
Jeffrey Lund
|
Piper Armstrong
|
Wilson Fearn
|
Stephen Cowley
|
Courtni Byun
|
Jordan Boyd-Graber
|
Kevin Seppi
Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics
Topic models are typically evaluated with respect to the global topic distributions that they generate, using metrics such as coherence, but without regard to local (token-level) topic assignments. Token-level assignments are important for downstream tasks such as classification. Even recent models, which aim to improve the quality of these token-level topic assignments, have been evaluated only with respect to global metrics. We propose a task designed to elicit human judgments of token-level topic assignments. We use a variety of topic model types and parameters and discover that global metrics agree poorly with human assignments. Since human evaluation is expensive we propose a variety of automated metrics to evaluate topic models at a local level. Finally, we correlate our proposed metrics with human judgments from the task on several datasets. We show that an evaluation based on the percent of topic switches correlates most strongly with human judgment of local topic quality. We suggest that this new metric, which we call consistency, be adopted alongside global metrics such as topic coherence when evaluating new topic models.
Search
Fix data
Co-authors
- Stephen Cowley 2
- Wilson Fearn 2
- Jeffrey Lund 2
- Kevin Seppi 2
- Jordan Boyd-Graber 1
- show all...