arrow
Return

Bipartite-Grammar-Aware Pretraining for XML-SQL Code Updating

delete2026-02-01
delete2
PRE
AI
Q
Qingyuan Liang
Z
Zeyu Sun *
Y
Yifan Zhao
Z
Zhihao Gong
G
Guoqing Wang
Y
Yizhou Chen
L
Lu Zhang *
G
Guangtai Liang *
Q
Qianxiang Wang
DOI:10.1145/3731752delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The eX tensible Markup Language (XML) is a file format widely used for data transmission in modern software development. In recent years, embedding SQL statements in XML files (i.e., XML-SQL) has become a popular way for developing applications with database access capability. Typically, XML-SQL code snippets demonstrate similar functionalities and structures, leading to repetitive programming work. Therefore, leveraging pretrained code models for automated code generation presents a promising way to alleviate duplicated efforts and enhance the efficiency of developing XML-SQL code. However, XML-SQL code has strong domainspecific characteristics that general pre-trained code models typically struggle to fully harness, thereby leading to limited overall performance of general pre-trained code models. In this article, we aim to address the challenge of handling this domain-specific knowledge. First, we propose a code updating task and construct the corresponding TwinXSQL dataset to better evaluate the model's code generation performance in the XML-SQL domain. Then, we leverage the common characteristics of XML-SQL and other programming languages (i.e., all programming languages impose grammar constraints on behavior) to design a bipartitegrammar-aware training framework (named BGA) for unsupervised pre-training, thereby improving the transfer of general-purpose code models to the XML-SQL domain. Specifically, we divide the XML-SQL code into two types of grammatical components: structure components and value components. During pretraining, we undertake three tasks, each designed to learn the internal information of these grammatical components and the relationships between them, enabling the pre-training process to better incorporate previously unlearned domain-specific knowledge of XML-SQL code. Our experimental results show that our trained model XSQLT5-base (220M) improves accuracy by 13.8% compared to the similarly sized CodeT5-base (220M). Additionally, our experiments reveal that ChatGPT, due to its inability to fully learn the XML-SQL domain knowledge, achieves a much lower generation accuracy even with few-shot samples compared to our XSQLT5-base (220M) model.
Keywords:
XML
Pre-training
Large Language Models
Code Generation
Software Evolution

Journal

A
ACM Transactions on Software Engineering and Methodology
IF:
6.2
Papers:
1.2K
Citations:
3.4K

Organization

I
institute of software, cas
Scholars:
445
Papers: 387
Citations: 0
P
peking university
Scholars:
11.7W
Papers: 8.7W
Citations: 146
C
chinese academy of sciences
Scholars:
56.0W
Papers: 44.8W
Citations: 704
researcher View more organizations