Return
Large-scale text-to-SQL generation with adversarial defense
DOI:10.1016/j.ins.2025.122942.png)
Abstract
En 中文
Large-scale Text-to-SQL models are vulnerable to perturbations in natural language query (NLQ) and database schemas. Most existing research has focused on adversarial attacks in input sequences, while neglecting defense. In this article, we argue that defense techniques are more critical, as they address diverse attacks. With this in mind, we posit that adversarial defense in large-scale Text-to-SQL poses a broader challenge than classical robustness. We also introduce two metrics to statistically evaluate defense performance. A framework for a certified robust method from an information theory perspective is proposed to address the new problem. One novel component is a regularizer (MIR), which uses active random masking to extract local features and maximize their mutual information with global features, ensuring theoretical robustness. Another new component is a Transformer-based Schema Linking (TSL) algorithm that enhances question schema alignment under adversarial settings. To support its supervised training, we propose Spider-SL, a new fine-grained alignment dataset derived from Spider. Our method is evaluated on five benchmarks encompassing 20 perturbation attacks. To the best of our knowledge, the results demonstrate that our model, using only 3B parameters, achieves state-of-the-art robustness and learning performance. This study suggests new research trends concerning the robustness of Text-to-SQL. Our code is available at: https://github.com/iliaohai/infosql.
Keywords:
Large-scale text-to-SQL
Robust adversarial defense
Mutual information
Schema linking
Semantic parsing
Journal
IF:
6.8
Papers:
540
Citations:
6.2W

