arrow
Return

Enhancing code smell classification with code refactoring-based data augmentation

delete2026-09-01
delete0
PRE
AI
A
Ali Nizam
M
Musa Aydin *
E
Ertuğrul İslamoğlu
S
Shaaban Sahmoud
S
Samil Bilal Ozaydin
DOI:10.1007/s10115-026-02859-2delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
One of the primary challenges in training deep neural networks is the requirement for a robust and diverse data source. To address this limitation, data augmentation techniques have emerged as a promising solution, enabling the expansion of training datasets without requiring additional data collection. This study investigates the effectiveness of code refactoring-based augmentation and the role of augmented data volume in improving code smell classification in deep learning systems. We conducted a comparative analysis of rule and large language model-based augmentation with a loss penalty, an algorithmic data balancing approach for neural networks. Additionally, a novel weighting strategy was developed to optimize the evaluation of augmentation volume distribution, reducing the computational overhead while maintaining analytical precision. We used pre-trained BERT, CodeBERT, and GraphCodeBERT models to generate code embeddings and evaluated their performance for model training. Our findings demonstrate that code refactoring-based augmentation using GraphCodeBERT and large language model enhances model performance, particularly in addressing class imbalances. The impact of data volume varies depending on the size of the classes, with the greatest improvement observed during the initial augmentation increments applied to underrepresented classes. Conversely, the influence of the loss penalty, BERT, and CodeBERT-based augmentation on overall performance was found to be negligible or, in some cases, detrimental. These results emphasize the potential of code refactoring-based augmentation to drive the development of more efficient data augmentation strategies, ultimately enabling better performance in code smell classification tasks and other deep learning applications.
Keywords:
Code embedding
Data augmentation
Code smell
Code refactoring
Loss penalty
Weight factor
Classification

Journal

Knowledge and Information Systems cover
Knowledge and Information Systems
IF:
3.1
Papers:
545
Citations:
5.2K

Organization

D
department of software engineering
Scholars:
136
Papers: 101
Citations: 0
D
Department of Computer Engineering
Scholars:
305
Papers: 166
Citations: 0
researcher View more organizations
Cited Papers

Cited Papers

Actionable code smell identification with fusion learning of metrics and semantics
err2024-09-01
err0
PREAI
errDongjin Yu; Quanxin Yang; Xin Chen; Jie Chen; Sixuan Wang; Yihang Xu
errShare
errSave
errShare
errSave
err
IF0
err
err0
PREAI
err
errShare
errSave
Code smells and refactoring: A tertiary systematic review of challenges and observations
err2020-09-01
err102
errOAAI
errLacerda, Guilherme; Petrillo, Fabio; Pimenta, Marcelo; Gueheneuc, Yann Gael
errShare
errSave
Nuanced Code Clone Detection Through LLM-Based Code Revision and AST Graph Modeling
err2025-01-01
err0
PREAI
errLi,Chunguang; Konpang,Jessada; Sirikham,Adisorn; Wang,Yan
errShare
errSave
Optimizing Pre-Trained Code Embeddings With Triplet Loss for Code Smell Detection
err2025-01-01
err0
errOAAI
errNizam, Ali; Islamoglu, Ertugrul; Kerem Adali, Omer; Aydin, Musa
errShare
errSave
Code smell detection by deep direct-learning and transfer-learning?
err2021-06-01
err51
errOAAI
errSharma, Tushar; Efstathiou, Vasiliki; Louridas, Panos; Spinellis, Diomidis
errShare
errSave
A Systematic Literature Review on Bad Smells-5 W's: Which, When, What, Who, Where
err2021-01-01
err58
PREAI
errde Paulo Sobrinho, Elder Vicente; De Lucia, Andrea; Maia, Marcelo de Almeida
errShare
errSave
researcher View more