Skip to main navigation Skip to search Skip to main content

ConDA: state-based data augmentation for context-dependent text-to-SQL

  • Dingzirui Wang
  • , Longxu Dou
  • , Wanxiang Che*
  • , Jiaqi Wang
  • , Jinbo Liu
  • , Lixin Li
  • , Jingan Shang
  • , Lei Tao
  • , Jie Zhang
  • , Cong Fu
  • , Xuri Song
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • State Grid Corporation of China
  • Tianjin Electric Power Corporation

Research output: Contribution to journalArticlepeer-review

Abstract

The context-dependent text-to-SQL task has profound real-world implications, as it facilitates users in extracting knowledge from vast databases, which allows users to acquire the information interactively for better accuracy. Unfortunately, current models struggle to address this task effectively due to the scarcity of data led by the high annotation overhead. The most straightforward method for addressing this problem is data augmentation, which aims at scaling up the parsing corpus. However, the naive methods suffer from the low diversity of the augmented data. To address this limitation, we propose the state-based CONtext-dependent text-to-SQL Data Augmentation (ConDA), which generate and filter augmented data based on the dialogue state, which has higher diversity. Experimental results show that ConDA yields performance improvement on all experimental datasets with an average boosting of 1.6%, proving the effectiveness of our method.

Original languageEnglish
Pages (from-to)3157-3168
Number of pages12
JournalInternational Journal of Machine Learning and Cybernetics
Volume15
Issue number8
DOIs
StatePublished - Aug 2024

Keywords

  • Context-dependent text-to-SQL
  • Data augmentation
  • Natural language processing
  • Semantic parsing

Fingerprint

Dive into the research topics of 'ConDA: state-based data augmentation for context-dependent text-to-SQL'. Together they form a unique fingerprint.

Cite this