Sumotosima : A Framework and Dataset for Classifying and Summarizing Otoscopic Images

Eram Anwarul Khan; Anas Anwarul Haq Khan

Sumotosima : A Framework and Dataset for Classifying and Summarizing Otoscopic Images

Eram Anwarul Khan, Anas Anwarul Haq Khan

Abstract

Otoscopy is a diagnostic procedure to examine the ear canal and eardrum using an otoscope. It identifies conditions like infections, foreign bodies, eardrum perforations, and ear abnormalities. We propose a novel resource-efficient deep learning and transformer-based framework, Sumotosima (Summarizer for Otoscopic Images), which provides an end-to-end pipeline for classification followed by summarization. Our framework utilizes a combination of triplet and cross-entropy losses. Additionally, we use Knowledge Enhanced Multimodal BART, where the input is fused textual and image embeddings. The objective is to deliver summaries that are well-suited for patients, ensuring clarity and efficiency in understanding otoscopic images. Given the lack of existing datasets, we have curated our own OCASD (Otoscopy Classification And Summary Dataset), which includes 500 images with 5 unique categories, annotated with their class and summaries by otolaryngologists. Sumotosima achieved a result of 98.03%, which is 7.00%, 3.10%, and 3.01% higher than K-Nearest Neighbors, Random Forest, and Support Vector Machines, respectively, in classification tasks. For summarization, Sumotosima outperformed GPT-4o and LLaVA by 88.53% and 107.57% in ROUGE scores, respectively. We have made our code and dataset publicly available at https://github.com/anas2908/Sumotosima

Anthology ID:: 2024.icon-1.1
Volume:: Proceedings of the 21st International Conference on Natural Language Processing (ICON)
Month:: December
Year:: 2024
Address:: AU-KBC Research Centre, Chennai, India
Editors:: Sobha Lalitha Devi, Karunesh Arora
Venue:: ICON
SIG:
Publisher:: NLP Association of India (NLPAI)
Note:
Pages:: 1–11
Language:
URL:: https://aclanthology.org/2024.icon-1.1/
DOI:
Bibkey:
Cite (ACL):: Eram Anwarul Khan and Anas Anwarul Haq Khan. 2024. Sumotosima : A Framework and Dataset for Classifying and Summarizing Otoscopic Images. In Proceedings of the 21st International Conference on Natural Language Processing (ICON), pages 1–11, AU-KBC Research Centre, Chennai, India. NLP Association of India (NLPAI).
Cite (Informal):: Sumotosima : A Framework and Dataset for Classifying and Summarizing Otoscopic Images (Khan & Khan, ICON 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.icon-1.1.pdf

PDF Cite Search Fix data