Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL

Yifei Shen; Yilun Zhao; Justice Ou; Tinglin Huang; Arman Cohan

Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL

Yifei Shen, Yilun Zhao, Justice Ou, Tinglin Huang, Arman Cohan

Abstract

Real-world clinical text-to-SQL requires reasoning over heterogeneous EHR tables, temporal windows, and patient-similarity cohorts to produce executable queries. We introduce ClinSQL, a benchmark of 633 expert-annotated tasks on MIMIC-IV v3.1 that demands multi-table joins, clinically meaningful filters, and executable SQL. Solving ClinSQL entails navigating schema metadata and clinical coding systems, handling long contexts, and composing multi-step queries beyond traditional text-to-SQL. We evaluate 20 proprietary and open-source models under Chain-of-Thought self-refinement and use rubric-based SQL analysis with execution checks that prioritize critical clinical requirements. Despite recent advances, performance remains far from clinical reliability: on the test set, GPT-5-mini attains 74.7% execution score, DeepSeek-R1 leads open-source at 69.2% and Gemini-2.5-Pro drops from 85.5% on Easy to 67.2% on Hard. Progress on ClinSQL marks tangible advances toward clinically reliable text-to-SQL for real-world EHR analytics.

Anthology ID:: 2026.eacl-long.64
Volume:: Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: March
Year:: 2026
Address:: Rabat, Morocco
Editors:: Vera Demberg, Kentaro Inui, Lluís Marquez
Venue:: EACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 1367–1412
Language:
URL:: https://aclanthology.org/2026.eacl-long.64/
DOI:
Bibkey:
Cite (ACL):: Yifei Shen, Yilun Zhao, Justice Ou, Tinglin Huang, and Arman Cohan. 2026. Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1367–1412, Rabat, Morocco. Association for Computational Linguistics.
Cite (Informal):: Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL (Shen et al., EACL 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.eacl-long.64.pdf
Checklist:: 2026.eacl-long.64.checklist.pdf

PDF Cite Search Checklist Fix data