arrow
Return

Improving Source Code Pre-Training via Type-Specific Masking

delete2025-02-20
delete0
delete
OA
AI
W
Wentao Zou
Q
Qi Li
李传艺 (Chuanyi Li)
葛季栋 (Jidong Ge)
陈翔 cover
陈翔 (Xiang Chen)
L
LiGuo Huang
B
Bin Luo
DOI:10.1145/3699599delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The Masked Language Modeling (MLM) task is widely recognized as one of the most effective pre-training tasks and currently derives many variants in the Software Engineering (SE) field. However, most of these variants mainly focus on code representation without distinguishing between different code token types, while some focus on a specific type, such as code identifiers. Indeed, various code token types exist, and there is no evidence that only identifiers can improve PTMs. Thus, to improve PTMs through different types, we conducted an extensive study to evaluate how different type-specific masking tasks can affect PTMs. First, we extract five code token types, convert them into type-specific masking tasks, and generate their combinations. Second, we pre-train CodeBERT and PLBART using combinations and fine-tuned them on four SE downstream tasks. Experimental results show that type-specific masking tasks can enhance CodeBERT and PLBART on all downstream tasks. Furthermore, we discuss topics related to low-resource datasets, conflicting PTMs that original pre-training tasks conflict with our methods, the cost and performance of our methods, factors that impact the performance of our methods, and applying our methods on state-of-the-art PTMs. These discussions comprehensively analyze the strengths and weaknesses of different type-specific masking tasks. CCS Concepts: center dot Software and its engineering- Software creation and management; center dot Computing methodologies- Artificial intelligence;
Keywords:
code representation learning
pre-trained model
pre-training tasks
masked language modeling

Journal

A
ACM Transactions on Software Engineering and Methodology
IF:
6.2
Papers:
1.2K
Citations:
3.4K

Organization

No organization information available