arrow
Return

A survey on code smells detection using machine learning techniques

delete2026-06-01
delete0
PRE
AI
M
Mesbah, Djamel *
N
Nour El Madhoun
A
Al Agha, Khaldoun
A
Anis Zouaoui
DOI:10.1016/j.infsof.2026.108242delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Code smells are widely used indicators of deteriorating design and maintainability and their detection has evolved from heuristic and metric-based rules to learning-based approaches that operate on structural and textual views of source code. Despite this evolution, existing surveys capture only parts of the landscape and provide limited insight into how modern Machine Learning and Deep Learning techniques, code representations and datasets interact in practice. This study offers a comprehensive synthesis of code smell detection research from 2013 to March 2025, covering heuristic, evolutionary, traditional Machine Learning and Deep Learning methods. We analyze detection techniques, the representations of code on which they rely, the datasets and benchmarks used for evaluation and the extent to which prior work supports reproducibility through available tools, code and replication packages. Building on this review, we carry out a comparative study on the MLCQ dataset in which 17 models from three families (classical ML on object-oriented metrics, sequencebased DL on token sequences and graph neural networks on abstract syntax trees) are evaluated on four smells (Feature Envy, Long Method, Blob and Data Class) under a unified pipeline. The results show that no single family dominates: GNNs excel on class-level smells (Blob, Data Class), sequence models lead on Long Method and Feature Envy remains an open challenge for all families due to the lack of inter-class context in current representations. The survey consolidates current knowledge on code smell detection and identifies open challenges related to dataset coverage, annotation quality and reproducibility, while the comparative study provides a transparent baseline that can inform and be extended by future evaluations.
Keywords:
Code smells
Static analysis
Machine learning
Deep learning
Software quality

Journal

Information and Software Technology cover
Information and Software Technology
IF:
4.3
Papers:
3.8K
Citations:
7.7K

Organization

U
Université Paris Saclay
Scholars:
430
Papers: 196
Citations: 0
C
Centre National de la Recherche Scientifique
Scholars:
1.8K
Papers: 803
Citations: 0
S
Sorbonne Universite
Scholars:
413
Papers: 183
Citations: 0
researcher View more organizations
Cited Papers

Cited Papers

No cited papers available