TY - GEN
T1 - TS-SQL
T2 - 30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025
AU - Xu, Wenbo
AU - Zhu, Haifeng
AU - Yan, Liang
AU - Liu, Chuanyi
AU - Han, Peiyi
AU - Duan, Shaoming
AU - Pan, Jeff Z.
N1 - Publisher Copyright:
©2025 Association for Computational Linguistics.
PY - 2025
Y1 - 2025
N2 - Large Language Model (LLM)-based self-refinement has advanced Text-to-SQL, but it struggles with SQL semantic errors, such as omitted conditions and misinterpreted requirements. This is because self-refinement depends on LLMs’ semantic understanding of questions, a process prone to hallucination-induced biases, leading to uncorrectable errors. To solve this problem, we propose Test-driven Self-refinement for Text-to-SQL (TS-SQL). It leverages a collaborative LLM agent framework to automatically synthesize high-quality test cases, including test data and test code. The test cases are further employed to provide execution feedback for LLM self-refinement towards SQL semantic errors. Rigorous evaluation shows the superiority of TS-SQL: for BIRD-dev, TS-SQL improves at least 6% over existing SQL self-refinement methods; for Spider-dev, TS-SQL identifies and corrects 131 gold SQL errors, exposing system flaws in benchmark rigor.
AB - Large Language Model (LLM)-based self-refinement has advanced Text-to-SQL, but it struggles with SQL semantic errors, such as omitted conditions and misinterpreted requirements. This is because self-refinement depends on LLMs’ semantic understanding of questions, a process prone to hallucination-induced biases, leading to uncorrectable errors. To solve this problem, we propose Test-driven Self-refinement for Text-to-SQL (TS-SQL). It leverages a collaborative LLM agent framework to automatically synthesize high-quality test cases, including test data and test code. The test cases are further employed to provide execution feedback for LLM self-refinement towards SQL semantic errors. Rigorous evaluation shows the superiority of TS-SQL: for BIRD-dev, TS-SQL improves at least 6% over existing SQL self-refinement methods; for Spider-dev, TS-SQL identifies and corrects 131 gold SQL errors, exposing system flaws in benchmark rigor.
UR - https://www.scopus.com/pages/publications/105028985362
U2 - 10.18653/v1/2025.findings-emnlp.156
DO - 10.18653/v1/2025.findings-emnlp.156
M3 - 会议稿件
AN - SCOPUS:105028985362
T3 - EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2025
SP - 2864
EP - 2889
BT - EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2025
A2 - Christodoulopoulos, Christos
A2 - Chakraborty, Tanmoy
A2 - Rose, Carolyn
A2 - Peng, Violet
PB - Association for Computational Linguistics (ACL)
Y2 - 4 November 2025 through 9 November 2025
ER -