KU-DMIS at EHRSQL 2024 : Generating SQL query via question templatization in EHR

Hajung Kim; Chanhwi Kim; Hoonick Lee; Kyochul Jang; Jiwoo Lee; Kyungjae Lee; Gangwoo Kim; Jaewoo Kang

doi:10.18653/v1/2024.clinicalnlp-1.64

KU-DMIS at EHRSQL 2024 : Generating SQL query via question templatization in EHR

Hajung Kim, Chanhwi Kim, Hoonick Lee, Kyochul Jang, Jiwoo Lee, Kyungjae Lee, Gangwoo Kim, Jaewoo Kang

Abstract

Transforming natural language questions into SQL queries is crucial for precise data retrieval from electronic health record (EHR) databases. A significant challenge in this process is detecting and rejecting unanswerable questions that request information outside the database’s scope or exceed the system’s capabilities. In this paper, we introduce a novel text-to-SQL framework that focuses on standardizing the structure of questions into a templated format. Our framework begins by fine-tuning GPT-3.5-turbo, a powerful large language model (LLM), with detailed prompts involving the table schemas of the EHR database system. Our approach shows promising results on the EHRSQL-2024 benchmark dataset, part of the ClinicalNLP shared task. Although fine-tuning GPT achieves third place on the development set, it struggled with the diverse questions in the test set. With our framework, we improve our system’s adaptability and achieve fourth position in the official leaderboard of the EHRSQL-2024 challenge.

Anthology ID:: 2024.clinicalnlp-1.64
Original:: 2024.clinicalnlp-1.64v1
Version 2:: 2024.clinicalnlp-1.64v2
Volume:: Proceedings of the 6th Clinical Natural Language Processing Workshop
Month:: June
Year:: 2024
Address:: Mexico City, Mexico
Editors:: Tristan Naumann, Asma Ben Abacha, Steven Bethard, Kirk Roberts, Danielle Bitterman
Venues:: ClinicalNLP | WS
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 672–686
Language:
URL:: https://aclanthology.org/2024.clinicalnlp-1.64/
DOI:: 10.18653/v1/2024.clinicalnlp-1.64
Bibkey:
Cite (ACL):: Hajung Kim, Chanhwi Kim, Hoonick Lee, Kyochul Jang, Jiwoo Lee, Kyungjae Lee, Gangwoo Kim, and Jaewoo Kang. 2024. KU-DMIS at EHRSQL 2024 : Generating SQL query via question templatization in EHR. In Proceedings of the 6th Clinical Natural Language Processing Workshop, pages 672–686, Mexico City, Mexico. Association for Computational Linguistics.
Cite (Informal):: KU-DMIS at EHRSQL 2024 : Generating SQL query via question templatization in EHR (Kim et al., ClinicalNLP 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.clinicalnlp-1.64.pdf

PDF (v2) PDF (v1) Cite Search Fix data