Bhavya Bhavya
2024
AnaDE1.0: A Novel Data Set for Benchmarking Analogy Detection and Extraction
Bhavya Bhavya
|
Shradha Sehgal
|
Jinjun Xiong
|
ChengXiang Zhai
Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)
Textual analogies that make comparisons between two concepts are often used for explaining complex ideas, creative writing, and scientific discovery. In this paper, we propose and study a new task, called Analogy Detection and Extraction (AnaDE), which includes three synergistic sub-tasks: 1) detecting documents containing analogies, 2) extracting text segments that make up the analogy, and 3) identifying the (source and target) concepts being compared. To facilitate the study of this new task, we create a benchmark dataset by scraping Metamia.com and investigate the performances of state-of-the-art models on all sub-tasks to establish the first-generation benchmark results for this new task. We find that the Longformer model achieves the best performance on all the three sub-tasks demonstrating its effectiveness for handling long texts. Moreover, smaller models fine-tuned on our dataset perform better than non-finetuned ChatGPT, suggesting high task difficulty. Overall, the models achieve a high performance on documents detection suggesting that it could be used to develop applications like analogy search engines. Further, there is a large room for improvement on the segment and concept extraction tasks.
Long-Form Analogy Evaluation Challenge
Bhavya Bhavya
|
Chris Palaguachi
|
Yang Zhou
|
Suma Bhat
|
ChengXiang Zhai
Proceedings of the 17th International Natural Language Generation Conference: Generation Challenges
Given the practical applications of analogies, recent work has studied analogy generation to explain concepts. However, not all generated analogies are of high quality and it is unclear how to measure the quality of this new kind of generated text. To address this challenge, we propose a shared task on automatically evaluating the quality of generated analogies based on seven comprehensive criteria. For this, we will set up a leader board based on our dataset annotated with manual ratings along the seven criteria, and provide a baseline solution leveraging GPT-4. We hope that this task would advance the progress in development of new evaluation metrics and methods for analogy generation in natural language, particularly for education.
2022
Analogy Generation by Prompting Large Language Models: A Case Study of InstructGPT
Bhavya Bhavya
|
Jinjun Xiong
|
ChengXiang Zhai
Proceedings of the 15th International Conference on Natural Language Generation
Search
Fix data
Co-authors
- ChengXiang Zhai 3
- Jinjun Xiong 2
- Suma Bhat 1
- Chris Palaguachi 1
- Shradha Sehgal 1
- show all...