FLEX: Expert-level False-Less EXecution Metric for Text-to-SQL Benchmark

Heegyu Kim; Jeon Taeyang; SeungHwan Choi; Seungtaek Choi; Hyunsouk Cho

doi:10.18653/v1/2025.naacl-long.228

FLEX: Expert-level False-Less EXecution Metric for Text-to-SQL Benchmark

Heegyu Kim, Jeon Taeyang, SeungHwan Choi, Seungtaek Choi, Hyunsouk Cho

Abstract

Text-to-SQL systems have become crucial for translating natural language into SQL queries in various industries, enabling non-technical users to perform complex data operations. The need for accurate evaluation methods has increased as these systems have grown more sophisticated. However, the Execution Accuracy (EX), the most prevalent evaluation metric, still shows many false positives and negatives. Thus, this paper introduces **FLEX(False-Less EXecution)**, a novel approach to evaluating text-to-SQL systems using large language models (LLMs) to emulate human expert-level evaluation of SQL queries. Our metric improves agreement with human experts (from 62 to 87.04 in Cohen’s kappa) with comprehensive context and sophisticated criteria. Our extensive experiments yield several key insights: (1) Models’ performance increases by over 2.6 points on average, substantially affecting rankings on Spider and BIRD benchmarks; (2) The underestimation of models in EX primarily stems from annotation quality issues; and (3) Model performance on particularly challenging questions tends to be overestimated. This work contributes to a more accurate and nuanced evaluation of text-to-SQL systems, potentially reshaping our understanding of state-of-the-art performance in this field.

Anthology ID:: 2025.naacl-long.228
Volume:: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
Month:: April
Year:: 2025
Address:: Albuquerque, New Mexico
Editors:: Luis Chiruzzo, Alan Ritter, Lu Wang
Venue:: NAACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 4448–4475
Language:
URL:: https://aclanthology.org/2025.naacl-long.228/
DOI:: 10.18653/v1/2025.naacl-long.228
Bibkey:
Cite (ACL):: Heegyu Kim, Jeon Taeyang, SeungHwan Choi, Seungtaek Choi, and Hyunsouk Cho. 2025. FLEX: Expert-level False-Less EXecution Metric for Text-to-SQL Benchmark. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 4448–4475, Albuquerque, New Mexico. Association for Computational Linguistics.
Cite (Informal):: FLEX: Expert-level False-Less EXecution Metric for Text-to-SQL Benchmark (Kim et al., NAACL 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.naacl-long.228.pdf

PDF Cite Search Fix data